Downloads · 30 days
45
45% of all-time downloads
GrimSqueaker/ProtSent-V2-ESMC-300M
ProtSent-V2-ESMC-300M is a sentence similarity model from GrimSqueaker. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as mit.
Contrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
Downloads · 30 days
45
45% of all-time downloads
All-time downloads
100
Public
Parameters
333M
1.3 GB on disk
Likes
2
Trending 1
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
Contrastively fine-tuned ESM-C 300M producing fixed-length protein embeddings where biological similarity maps to embedding proximity. Intended for retrieval, clustering, and nearest-neighbour transfer.
This is the ProtSent V2 recipe applied to an ESM-C backbone.
Continued training from the Synthyra/ESMplusplus_small (ESMC-300M) weights.
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
model = SentenceTransformer("GrimSqueaker/ProtSent-V2-ESMC-300M", trust_remote_code=True)
emb = model.encode([
"MKTLLLTLVVVTIVCLDLGYT",
"MKTLLLTLVVVTIVCLDLGYN",
"AGWYRSPQEGLKPVDTFKDIV",
])
print(cos_sim(emb[0], emb[1:]))
Embeddings are mean-pooled over the final layer, dimension 960. No Matryoshka heads on this release (see Training configuration).
Four sources. Pfam/AFDB/STRING are decontaminated against benchmark test sets (see below); the DMS source is not — see Limitations.
| source | pairs used |
|---|---|
| Pfam families | 1,230,551 |
| AlphaFold DB (Foldseek clusters) | 27,105,179 |
| STRING-DB v12 PPI | 15,000,000 |
| ProteinGym DMS (CoSENT target) | 1,000,000 |
| total | 44,335,730 |
Pfam/AFDB pairs are drawn with the disjoint-pair sampler at up to 10 pairs per cluster
(k=10), which spends the pair budget on distinct sequence pairs rather than repeating
members — covers up to 90 distinct sequences per cluster instead of 10. STRING is a fixed
15M-pair subsample (seed 42).
Pfam, AlphaFold DB and STRING were searched against the benchmark test sequences with
MMseqs2 easy-search (40% identity, 80% coverage, --cov-mode 1) and matching sequences
removed before training — the same protsent-data-dc40 corpus used by the ProtSent-V2.5
ESM-2 releases.
| setting | value |
|---|---|
| backbone | ESM-C 300M (Synthyra/ESMplusplus_small, 960 hidden) |
| loss | CachedMultipleNegativesRankingLoss (Pfam/AFDB/STRING) + CoSENT (DMS) |
| contrastive batch | 2048 per device |
| gather across devices | off |
| multi-dataset sampler | proportional, world-size-aligned (see below) |
| GOR weight | 0 (off) |
| Matryoshka | off |
| max sequence length | 512 |
| optimiser | AdamW, LR 1e-4, warmup 300 steps |
| precision / attention | bf16, flash-attention-2 (FA3 unavailable on this hardware) |
| hardware | 3x NVIDIA B300 |
| steps | 7,216 (one epoch) |
| gradient-cache mini-batch | 128 |
Training code: github.com/oriel9p/ProtSent,
train_esmc_300m_v2.sh.
23-task suite, -e test, seed 42, both kNN and linear-probe transfer. Full per-task CSVs:
results/benchmarks/ism/protsent_esmc_300m_v2_{knn,linear}/ in the repo above.
| method | R@1 | R@10 | MAP |
|---|---|---|---|
| ESM-C 300M (this model's backbone, untrained) | 0.3709 | 0.5794 | 0.2212 |
| ISM-C 300M (structure-distilled ESM-C 300M) | 0.4300 | 0.6592 | 0.2733 |
| HMMER (phmmer, filters off) | 0.7525 | 0.8978 | 0.6067 |
| ProtSent-V2 ESM-C 300M | 0.7974 | 0.9539 | 0.7692 |
Win/tie/loss, tie tolerance 0.005:
| probe | vs ESM-C 300M (this backbone, untrained) | vs ISM-C 300M |
|---|---|---|
| kNN | 16W / 0T / 7L | 14W / 1T / 8L |
| linear | 6W / 3T / 14L | 6W / 4T / 13L |