Downloads · 30 days
99
31% of all-time downloads
sabsab129/MiniLM-searchkeys
MiniLM-searchkeys is a sentence similarity model from sabsab129. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as cc.
MiniLM-searchkeys is a sentence-transformers encoder fine-tuned for multi-domain keyphrase ranking as part of SearchKeys, a retrieval-augmented encoder-only alternative to seq2seq keyphrase prediction. It maps documen…
Downloads · 30 days
99
31% of all-time downloads
All-time downloads
324
Public
Parameters
33.4M
401 MB on disk
Likes
0
Public
Click a slice to open those files.
.pt267 MB · 67%
From the Hugging Face model README
MiniLM-searchkeys is a sentence-transformers encoder fine-tuned for multi-domain keyphrase ranking as part of SearchKeys, a retrieval-augmented encoder-only alternative to seq2seq keyphrase prediction. It maps documents and candidate keyphrases into a shared 384-dimensional space, where cosine similarity reflects keyphrase relevance — including for absent keyphrases (phrases that do not literally appear in the source text).
This model is the fine-tuned encoder described in:
Saber Zahhar, Nédra Mellouli, Christophe Rodrigues, Nicolas Travers. Multi-Domain Keyphrase Prediction via Retrieval-Augmented Ranking: A Resource-Efficient Alternative to Seq2Seq Generation. DKE 2026.
Code and full pipeline: github.com/saberzahhar/dke2026kp
Instead of generating keyphrases token-by-token, SearchKeys retrieves and ranks:
Because candidates come from a real annotated corpus rather than a vocabulary distribution, this approach is able to surface relevant keyphrases that are absent from the source document — something generative decoders struggle with.
This model starts from the pretrained sentence-transformers/all-MiniLM-L12-v2 checkpoint and is fine-tuned with a contrastive, multi-label objective tailored to keyphrase ranking.
Fine-tuning combines three multi-domain keyphrase datasets, used jointly (not domain-specialized):
| Dataset | Domain | Train docs |
|---|---|---|
| kp20k | Computer science (ACM DL, ScienceDirect, Wiley, etc.) | 530.8k |
| kpbiomed | Biomedical (PubMed) | 500k |
| kptimes | News (NYTimes / Japan Times) | 259.9k |
This model is intended to be used as the ranking encoder in a retrieval-augmented keyphrase prediction pipeline, not as a general-purpose sentence embedder. Given a document and a pool of candidate keyphrases (e.g. pooled via BM25 retrieval from an annotated corpus), it produces embeddings whose cosine similarity is a strong relevance signal for both present and absent keyphrases.
It is best paired with:
all-MiniLM-L12-v2 and mxbai-embed-large-v1 on candidate recall);By default, input text longer than 512 word pieces is truncated.
pip install -U sentence-transformers
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
model = SentenceTransformer("sabsab129/MiniLM-searchkeys")
document = "A feedback vertex set of 2-degenerate graphs..."
candidates = [
"feedback vertex set",
"decycling set",
"2-degenerate graphs",
"rank-width",
"fixed-parameter algorithm",
]
doc_embedding = model.encode(document)
candidate_embeddings = model.encode(candidates)
scores = cos_sim(doc_embedding, candidate_embeddings)[0]
ranked = sorted(zip(candidates, scores.tolist()), key=lambda x: -x[1])
print(ranked)
If you use this model, please cite the paper:
@inproceedings{zahhar2026searchkeys,
title = {Multi-Domain Keyphrase Prediction via Retrieval-Augmented Ranking: A Resource-Efficient Alternative to Seq2Seq Generation},
author = {Zahhar, Saber and Mellouli, N{\'e}dra and Rodrigues, Christophe and Travers, Nicolas},
booktitle = {DKE 2026},
year = {2026}
}
This model is fine-tuned from sentence-transformers/all-MiniLM-L12-v2, originally developed by the Sentence-Transformers team during the Hugging Face JAX/Flax Community Week, based on microsoft/MiniLM-L12-H384-uncased.