Downloads · 30 days
177
24% of all-time downloads
BatuhanECB/FinModernBERT-embed-large-v1
FinModernBERT-embed-large-v1 is a sentence similarity model from BatuhanECB. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
A 395M-parameter finance text-embedding model that beats the 7B Fin-E5 on FinMTEB Summarization (+0.109) and STS (+0.010) at 1/18th the size — built with ≈120 GPU-hours (~5 days) on a single 2× RTX 3090 workstation, t…
Downloads · 30 days
177
24% of all-time downloads
All-time downloads
742
Public
Parameters
395M
1.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.6 GB · 100%
From the Hugging Face model README
A 395M-parameter finance text-embedding model that beats the 7B Fin-E5 on FinMTEB Summarization (+0.109) and STS (+0.010) at 1/18th the size — built with ≈120 GPU-hours (~5 days) on a single 2× RTX 3090 workstation, through a fully decontaminated four-step pipeline: measured baseline → 5.65B-token domain MLM (DAPT, ~85 of those GPU-hours) → 270k synthetic contrastive pairs → multi-task InfoNCE (~8 h) → WiSE-FT weight interpolation (0 training).
Now on the official FinMTEB leaderboard (18 EN models): STS #2 (the top embedding model — only a lexical BOW baseline scores higher) and Summarization #2 (behind only voyage-3-large), with the single best FINDsum score on the board. See §2 for the full placement.
This card documents the entire journey: the goal, every training stage, every measured number, the failure that almost sank the release, and the trick that fixed it.
Build the strongest defensible finance embedder on a single encoder of ModernBERT-large size, evaluated on the full English FinMTEB benchmark (35 tasks, 7 task types), against the published state of the art — Fin-E5 (FinMTEB paper), a finance-adapted e5-mistral-7B. Target: win STS and Summarization outright, make Retrieval competitive.
"Defensible" means: every training stage runs behind a hard decontamination gate against all FinMTEB English eval sets (word-shingle overlap index; eval-source datasets and contaminated lineages like FiQA/ConvFinQA excluded outright), and the eval harness is the official FinMTEB fork end to end.
| Task type | n | This model (395M) | Fin-E5 (7B, published) | e5-mistral base (7B) | Δ vs Fin-E5 |
|---|---|---|---|---|---|
| Summarization | 3 | 0.588 | 0.480 | 0.528 | +0.109 👑 |
| STS | 2 | 0.444 | 0.434 | 0.380 | +0.010 👑 |
| Reranking | 3 | 0.961 | 0.990 | 0.988 | −0.028 |
| Clustering | 6 | 0.513 | 0.565 | 0.578 | −0.052 |
| Classification | 8 | 0.623 | 0.757 | 0.645 | −0.133 |
| PairClassification | 3 | 0.617 | 0.801 | 0.739 | −0.184 |
| Retrieval | 10 | 0.502 | 0.711 | 0.675 | −0.208 |
| Overall (7-type mean) | 35 | 0.607 | 0.677 | 0.648 | −0.070 |
Fin-E5 / e5-mistral rows are the published Table-1 numbers (weights are not public); this model was scored locally on the same FinMTEB fork with identical task mains (nDCG@10 retrieval, Spearman STS/Summ, MAP reranking, accuracy/v-measure/AP elsewhere).
The model is now listed on the official FinMTEB leaderboard (English board: 18 models, including voyage-3-large, OpenAI text-embedding-3-large/small, NV-Embed v2, e5-mistral-7B, gte-Qwen1.5-7B, bge-en-icl and Fin-E5). Position as of 2026-07-19, ranks computed from the leaderboard's own per-task scores by averaging tasks within each type:
| Task type (EN) | Score | Rank | Note |
|---|---|---|---|
| STS | 0.444 | #2 / 18 | #1 among all embedding models — only the lexical bag-of-words baseline (0.485) sits higher; ahead of Fin-E5 (0.434), voyage-3-large (0.415), NV-Embed v2 (0.374) |
| Summarization | 0.588 | #2 / 18 | behind only voyage-3-large (0.648); ahead of text-embedding-3-large (0.567) and Fin-E5 (0.480). FINDsum 0.745 is the single best score on the whole board (next: voyage 0.700) |
| Classification | 0.623 | #11 / 18 | |
| Reranking | 0.961 | #13 / 18 | |
| Clustering | 0.513 | #13 / 18 | |
| Retrieval | 0.502 | #14 / 18 | |
| PairClassification | 0.617 | #15 / 18 | |
| Overall (7-type mean) | 0.607 | #11 / 18 | −0.070 behind the #1 (Fin-E5 0.677) at 1/18th its size |
Only four models on the board achieve two or more top-2 task-type finishes: voyage-3-large and text-embedding-3-large (closed-source APIs), Fin-E5 (7B, weights not public) — and this 395M open-weights model. The leaderboard's numbers for this model match the locally-measured scores reported below exactly.
Everything starts from a measured baseline of the untrained base model against the finance SOTA, on a 15-task FinMTEB subset (2 STS + 3 Summarization + 10 Retrieval):
| task type | ModernBERT-large (untrained) | Fin-Retriever-base (110M, trained) | Fin-E5 (published) |
|---|---|---|---|
| STS | 0.450 | 0.294 | 0.434 |
| Summarization | 0.118 | 0.281 | 0.480 |
| Retrieval | 0.051 | 0.402 | 0.711 |
Two findings shaped the whole project: (a) untrained ModernBERT-large already matches Fin-E5 on STS — that lead must survive training; (b) Retrieval (0.051) and Summarization (0.118) are where the crown is won or lost.
Released separately as FinModernBERT-large-DAPT.
FinMTEB Summarization documents are huge (FNS averages ~290k chars ≈ 70k+ tokens). A zero-training experiment showed head-truncation was destroying the signal: chunk the document into ≤16 windows of ~506 body tokens, embed each, L2-normalize, mean-pool, re-normalize → Summarization 0.120 → 0.305 with no training at all (FNS alone: 0.223 → 0.707). This became the official doc-side embedding strategy.
intfloat/e5-large-v2, exact top-50 then filtered by sim(cand) ≤ 0.95 · sim(positive)
to avoid false negatives; ≥95% of anchors got 4 negatives.Pure contrastive training delivered Retrieval 0.569 and Summarization 0.625 — but STS collapsed 0.450 → 0.371, below the pre-registered release gate (≥ 0.434 = Fin-E5 parity). Both epoch checkpoints failed it.
Instead of retraining, the zero-cost fix: interpolate encoder weights back toward the
DAPT initialization (WiSE-FT): w = α·contrastive + (1−α)·DAPT, swept on the eval:
| α (contrastive share) | STS | Summarization | Retrieval | gate ≥0.434 |
|---|---|---|---|---|
| 1.00 (pure) | 0.371 | 0.625 | 0.569 | ✗ |
| 0.80 | 0.402 | — | — | ✗ |
| 0.65 → released | 0.444 | 0.588 | 0.502 | ✓ |
| 0.50 | 0.465 | 0.457 | 0.327 | ✓ (retrieval collapses) |
| 0.30 | 0.446 | — | — | ✓ |
| 0.00 (DAPT) | 0.446 | 0.305 | 0.056 | ✓ |
Two lessons: α=0.5 scores above both endpoints on STS (the classic WiSE-FT bump), and retrieval skill decays steeply toward the DAPT end — α=0.65 is the balance point that passes the gate while keeping the Summarization crown and most of the retrieval gain.
| Task | Type | Score |
|---|---|---|
| FINAL | STS | 0.5885 |
| FinSTS | STS | 0.3003 |
| FNS2022sum | Summarization | 0.8528 |
| FINDsum | Summarization | 0.7453 |
| Ectsum | Summarization | 0.1667 |
| Apple10KRetrieval | Retrieval | 0.8808 |
| TradeTheEventEncyclopediaRetrieval | Retrieval | 0.8129 |
| TradeTheEventNewsRetrieval | Retrieval | 0.7540 |
| FinanceBenchRetrieval | Retrieval | 0.5853 |
| USNewsRetrieval | Retrieval | 0.5369 |
| HC3Retrieval | Retrieval | 0.4216 |
| TheGoldmanEnRetrieval | Retrieval | 0.3891 |
| FiQA2018Retrieval | Retrieval | 0.2886 |
| TATQARetrieval | Retrieval | 0.1885 |
| FinQARetrieval | Retrieval | 0.1633 |
| FinFactReranking | Reranking | 0.9752 |
| HC3Reranking | Reranking | 0.9672 |
| FiQA2018Reranking | Reranking | 0.9412 |
| ESGClassification | Classification | 0.8144 |
| FinancialPhraseBankClassification | Classification | 0.7815 |
| FinancialFraudClassification | Classification | 0.6392 |
| FiQAClassification | Classification | 0.6074 |
| FinSentClassification | Classification | 0.5854 |
| FLSClassification | Classification | 0.5648 |
| SemEva2017Classification | Classification | 0.5618 |
| FOMCClassification | Classification | 0.4319 |
| PiiClustering | Clustering | 0.8619 |
| MInDS14EnClustering | Clustering | 0.8268 |
| WikiCompany2IndustryClustering | Clustering | 0.6831 |
| ComplaintsClustering | Clustering | 0.2791 |
| FinanceArxivS2SClustering | Clustering | 0.2146 |
| FinanceArxivP2PClustering | Clustering | 0.2134 |
| HeadlinePDDPairClassification | PairClassification | 0.6371 |
| HeadlineACPairClassification | PairClassification | 0.6073 |
| HeadlinePDUPairClassification | PairClassification | 0.6073 |
| Milestone | STS | Summarization | Retrieval |
|---|---|---|---|
| Stage 0: untrained ModernBERT-large | 0.450 | 0.118 | 0.051 |
| Stage 1: + 5.65B-token DAPT | 0.446 | 0.120 | 0.056 |
| + chunk+mean-pool eval strategy | 0.446 | 0.305 | 0.056 |
| Stage 3: + contrastive (pure) | 0.371 | 0.625 | 0.569 |
| + WiSE-FT α=0.65 (this model) | 0.444 | 0.588 | 0.502 |
| Fin-E5 7B (the bar) | 0.434 | 0.480 | 0.711 |
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BatuhanECB/FinModernBERT-embed-large-v1")
# Asymmetric retrieval: prefix queries and passages
queries = ["query: What drove the increase in operating expenses?"]
passages = [
"passage: Operating expenses rose 12% year-over-year, driven primarily by "
"increased headcount in R&D and higher cloud infrastructure costs.",
]
q = model.encode(queries, normalize_embeddings=True)
p = model.encode(passages, normalize_embeddings=True)
print(q @ p.T)
# Symmetric use (STS / clustering): prefix both sides with "passage: "
sents = ["passage: Net revenue increased 8%.", "passage: Sales grew by eight percent."]
emb = model.encode(sents, normalize_embeddings=True)
Prefixes matter — training baked literal query: / passage: prefixes in.
Long documents (>512 tokens): reproduce the eval numbers by chunking into ≤16 windows
of ~506 body tokens, embedding each, L2-normalizing, mean-pooling, re-normalizing.
Embedding dim 1024 · mean pooling · cosine similarity · max_seq 512 (native ModernBERT window is 8,192; 512 is the training configuration).
Planned next iteration: round-2 self-mining with this model, STS-preserving retrain (dropping the need for interpolation), synthetic table-QA pairs, finance-NLI headline pairs, and label-based contrastive for classification.
Model weights: Apache-2.0. The DAPT corpus and synthetic pair set are not released (source licensing does not permit redistribution); the encoder does not memorize/regurgitate corpus text in embedding use. Every training stage was decontaminated against all FinMTEB English eval sets before training.