Downloads · 30 days
0
sugiv/modernbert-us-stablecoin-encoder
modernbert-us-stablecoin-encoder is a sentence similarity model from sugiv. Use it when you need a score for how close two texts are. It is set up for transformers. The card lists the license as apache-2.0.
Production-ready encoder model for semantic search over US stablecoin regulatory documents.
Downloads · 30 days
0
Access
Public
Updated Mar 11, 2026
Repo size
9.2 MB
Likes
0
Public
Click a slice to open those files.
.safetensors9.2 MB · 72%
From the Hugging Face model README
Production-ready encoder model for semantic search over US stablecoin regulatory documents.
Fine-tuned from answerdotai/ModernBERT-base using LoRA on 10,260 synthetic query-document triplets.
| Metric | Score | vs Base Model |
|---|---|---|
| NDCG@10 | 0.9236 | +517% (5.2x better) 🔥 |
| MRR@10 | 0.8991 | +714% (7.1x better) 🔥 |
| Recall@10 | 0.9961 | +250% (2.5x better) 🔥 |
| Recall@100 | 1.0000 | Perfect - never misses docs |
pip install "transformers>=4.48" "peft>=0.14" "huggingface_hub>=0.27" torch
Important: ModernBERT requires
transformers >= 4.48. Older versions will fail withKeyError: 'modernbert'.
| Library | Minimum Version | Notes |
|---|---|---|
transformers | >= 4.48 | ModernBERT architecture support |
peft | >= 0.14 | Compatible hf_hub_download API |
huggingface_hub | >= 0.27 | No deprecated use_auth_token |
torch | >= 2.0 | CUDA support |
pip install "transformers>=4.48" "peft>=0.14" "huggingface_hub>=0.27" torch
Tokenizer: Load from the base model (answerdotai/ModernBERT-base), NOT from this adapter repo. The adapter repo stores LoRA weights only; the tokenizer lives with the base model.
UNEXPECTED keys on load: When loading AutoModel.from_pretrained("answerdotai/ModernBERT-base"), you may see warnings about UNEXPECTED keys (head.norm.weight, head.dense.weight, decoder.bias). These are the base model's MLM (masked language model) head weights that exist in the pretrained checkpoint but are not used by AutoModel (which loads only the encoder backbone). This is completely normal and safe to ignore.
from transformers import AutoModel, AutoTokenizer
from peft import PeftModel
import torch
# Load model
base_model = AutoModel.from_pretrained(
"answerdotai/ModernBERT-base",
trust_remote_code=True
)
model = PeftModel.from_pretrained(
base_model,
"sugiv/modernbert-us-stablecoin-encoder"
)
model.eval()
# IMPORTANT: Load tokenizer from BASE MODEL, not from adapter repo
tokenizer = AutoTokenizer.from_pretrained(
"answerdotai/ModernBERT-base",
trust_remote_code=True
)
def encode(text, max_length=512):
inputs = tokenizer(
text, padding=True, truncation=True,
max_length=max_length, return_tensors="pt"
)
with torch.no_grad():
outputs = model(**inputs)
embeddings = outputs.last_hidden_state.mean(dim=1)
embeddings = embeddings / embeddings.norm(dim=1, keepdim=True)
return embeddings
# Example
query = "What are reserve requirements for stablecoin issuers?"
query_emb = encode(query)
Trained to understand US stablecoin regulatory concepts:
Metrics Explained:
Validation: 10,260 queries × 38 docs = full corpus ranking
Apache 2.0 (following ModernBERT-base)
Status: 🟢 Production Ready
Updated: March 10, 2026
Version: 1.0.0