Downloads · 30 days
86
73% of all-time downloads
TensorVizion/roBERTa-nanobeir
roBERTa-nanobeir is a sentence similarity model from TensorVizion. Use it when you need a score for how close two texts are.
This is a roberta-base model I fine-tuned on the NanoBEIR dataset for semantic search and passage retrieval. It's not chasing the top of any leaderboard — it's a compact, fast-loading embedding model that runs happily…
Downloads · 30 days
86
73% of all-time downloads
All-time downloads
118
Public
Parameters
125M
499 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors499 MB · 99%
From the Hugging Face model README
This is a roberta-base model I fine-tuned on the NanoBEIR dataset for semantic search and passage retrieval. It's not chasing the top of any leaderboard — it's a compact, fast-loading embedding model that runs happily on modest hardware and still retrieves sensibly.
Most retrieval models on the Hub are enormous. Great results, sure — but they need serious GPUs just to load. I wanted something that:
NanoBEIR felt like the right training signal for that. It's a compact collection of retrieval tasks that mirrors the diversity of the full BEIR benchmark, so the model gets exposed to a bit of everything — questions, abstracts, web-style queries — without me needing a compute cluster to train it.
The easy way, with sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("YOUR_USERNAME/roberta-nanobeir")
queries = ["what causes the seasons"]
passages = [
"The seasons are caused by the tilt of the Earth's rotational axis relative to its orbit around the Sun.",
"To fix a leaky tap, first shut off the water supply under the sink.",
]
query_emb = model.encode(queries)
passage_emb = model.encode(passages)
scores = model.similarity(query_emb, passage_emb)
print(scores) # the first passage should win, comfortably
If you'd rather stay in plain transformers, mean pooling + normalisation works fine too:
import torch
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("YOUR_USERNAME/roberta-nanobeir")
model = AutoModel.from_pretrained("YOUR_USERNAME/roberta-nanobeir")
def embed(texts):
inputs = tokenizer(texts, padding=True, truncation=True,
max_length=512, return_tensors="pt")
with torch.no_grad():
out = model(**inputs)
mask = inputs["attention_mask"].unsqueeze(-1)
emb = (out.last_hidden_state * mask).sum(1) / mask.sum(1)
return torch.nn.functional.normalize(emb, p=2, dim=1)
Nothing exotic here — a pretty standard contrastive setup: | Batch size | 4 | | Learning rate | 2e-5 |
The idea is simple: pull the matching passage close to its query in embedding space, push everything else in the batch away. Repeat a few hundred thousand times and the model slowly learns what "relevant" means across a bunch of different domains.
A few things you should know before using this in anything real:
The weights are released under the MIT license (inherited from roberta-base). Note that NanoBEIR is assembled from several source corpora with their own licenses — if you redistribute training data or build something commercial on top, it's worth checking the constituent dataset licenses yourself.
If this model is useful to you, or if you find a task where it embarrasses itself, open a discussion on the model card — I'd genuinely like to know. Retrieval models improve through people reporting where they fail, not through vibes.