Downloads · 30 days
14.2K
16% of all-time downloads
dleemiller/ModernCE-base-sts
ModernCE-base-sts is a text classification model from dleemiller. Use it when you need a label for a piece of text. It is set up for sentence-transformers. The card lists the license as mit.
Cross encoders are high performing encoder models that compare two texts and output a 0-1 score. I've found the cross-encoders/roberta-large-stsb model to be very useful in creating evaluators for LLM outputs. They're…
Downloads · 30 days
14.2K
16% of all-time downloads
All-time downloads
86.8K
Public
Parameters
150M
1.6 GB on disk
Likes
11
Trending 1
Click a slice to open those files.
.onnx1 GB · 64%
From the Hugging Face model README
Cross encoders are high performing encoder models that compare two texts and output a 0-1 score.
I've found the cross-encoders/roberta-large-stsb model to be very useful in creating evaluators for LLM outputs.
They're simple to use, fast and very accurate.
Like many people, I was excited about the architecture and training uplift from the ModernBERT architecture (answerdotai/ModernBERT-base).
So I've applied it to the stsb cross encoder, which is a very handy model. Additionally, I've added
pretraining from a much larger semi-synthetic dataset dleemiller/wiki-sim that targets this kind of objective.
The inference performance efficiency, expanded context and simplicity make this a really nice platform as an evaluator model.
dleemiller/wiki-sim and fine-tuned on sentence-transformers/stsb.| Model | STS-B Test Pearson | STS-B Test Spearman | Context Length | Parameters | Speed |
|---|---|---|---|---|---|
dleemiller/ModernCE-large-sts | 0.9256 | 0.9215 | 8192 | 395M | Medium |
dleemiller/CrossGemma-sts-300m | 0.9175 | 0.9135 | 2048 | 303M | Medium |
dleemiller/ModernCE-base-sts | 0.9162 | 0.9122 | 8192 | 149M | Fast |
cross-encoder/stsb-roberta-large | 0.9147 | - | 512 | 355M | Slow |
dleemiller/EttinX-sts-m | 0.9143 | 0.9102 | 8192 | 149M | Fast |
dleemiller/NeoCE-sts | 0.9124 | 0.9087 | 4096 | 250M | Fast |
dleemiller/EttinX-sts-s | 0.9004 | 0.8926 | 8192 | 68M | Very Fast |
cross-encoder/stsb-distilroberta-base | 0.8792 | - | 512 | 82M | Fast |
dleemiller/EttinX-sts-xs | 0.8763 | 0.8689 | 8192 | 32M | Very Fast |
dleemiller/EttinX-sts-xxs | 0.8414 | 0.8311 | 8192 | 17M | Very Fast |
dleemiller/sts-bert-hash-nano | 0.7904 | 0.7743 | 8192 | 0.97M | Very Fast |
dleemiller/sts-bert-hash-pico | 0.7595 | 0.7474 | 8192 | 0.45M | Very Fast |
To use ModernCE for semantic similarity tasks, you can load the model with the Hugging Face sentence-transformers library:
from sentence_transformers import CrossEncoder
# Load ModernCE model
model = CrossEncoder("dleemiller/ModernCE-base-sts")
# Predict similarity scores for sentence pairs
sentence_pairs = [
("It's a wonderful day outside.", "It's so sunny today!"),
("It's a wonderful day outside.", "He drove to work earlier."),
]
scores = model.predict(sentence_pairs)
print(scores) # Outputs: array([0.9184, 0.0123], dtype=float32)
The model returns similarity scores in the range [0, 1], where higher scores indicate stronger semantic similarity.
The model was pretrained on the pair-score-sampled subset of the dleemiller/wiki-sim dataset. This dataset provides diverse sentence pairs with semantic similarity scores, helping the model build a robust understanding of relationships between sentences.
cross-encoder/stsb-roberta-large.Fine-tuning was performed on the sentence-transformers/stsb dataset.
The model achieved the following test set performance after fine-tuning:
dleemiller/wiki-sim (pair-score-sampled)sentence-transformers/stsbThanks to the AnswerAI team for providing the ModernBERT models, and the Sentence Transformers team for their leadership in transformer encoder models.
If you use this model in your research, please cite:
@misc{moderncestsb2025,
author = {Miller, D. Lee},
title = {ModernCE STS: An STS cross encoder model},
year = {2025},
publisher = {Hugging Face Hub},
url = {https://huggingface.co/dleemiller/ModernCE-base-sts},
}
This model is licensed under the MIT License.