Downloads · 30 days
94
100% of all-time downloads
seamon67/Supertron-2-Reranker-2B-GGUF
Supertron-2-Reranker-2B-GGUF is a text ranking model from seamon67. Use it for the text ranking task on the model card, and read the license before you ship it in a product. It is set up for llama.cpp. The card lists the license as apache-2.0.
This model was converted to GGUF format from Surpem/Supertron2-Reranker-2B using a modified version of llama.cpp (release b10269). Refer to the original model card for more details on the model.
Downloads · 30 days
94
100% of all-time downloads
All-time downloads
94
Public
Repo size
2.7 GB
Likes
0
Public
Click a slice to open those files.
.gguf2.7 GB · 100%
From the Hugging Face model README
This model was converted to GGUF format from Surpem/Supertron2-Reranker-2B using a modified version of llama.cpp (release b10269). Refer to the original model card for more details on the model.
This is a working GGUF. It will work on any modern version of llama.cpp.
Most community GGUFs of Supertron2-Reranker-2B produce garbage scores because of 3 issues:
sigmoid(yes_logit - no_logit).This GGUF fixes all 3 problems.
Supertron2-Reranker-2B is a reranking model built on top of Qwen/Qwen3-VL-Reranker-2B. It is designed to score query-document pairs for retrieval pipelines, search systems, and RAG applications where a stronger second-stage ranker is useful.
Supertron2-Reranker-2B can compare a user query against candidate passages and assign relevance scores. It is intended as a second-stage reranker after a faster retriever has already selected candidate documents.
The model can help improve retrieval-augmented generation by pushing more relevant documents toward the top of the context window before answer generation.
Supertron2-Reranker-2B is useful for matching questions to passages, snippets, help-center articles, documentation chunks, and other text candidates.
The model is prompted for relevance scoring, making it suitable for natural language search tasks where query intent matters.
from sentence_transformers import CrossEncoder
model_id = "Surpem/Supertron2-Reranker-2B"
model = CrossEncoder(model_id)
pairs = [
("What is the capital of France?", "Paris is the capital and largest city of France."),
("What is the capital of France?", "Mars is often called the red planet."),
]
scores = model.predict(pairs)
print(scores)
Example reranking:
query = "How do I reset my password?"
documents = [
"Use the account recovery page to reset your password.",
"Our refund policy allows returns within 30 days.",
"Two-factor authentication adds extra login security.",
]
results = model.rank(query, documents)
print(results)
| Precision | Min VRAM | Recommended |
|---|---|---|
| bfloat16 | 6 GB | 10 GB+ |
| 4-bit quantized | 3 GB | 6 GB+ |
For larger batches or long documents, use more VRAM or reduce the batch size/max sequence length.
Supertron2-Reranker-2B is intended for:
It is not intended to be used as a standalone chat model.
@misc{surpem2026supertron2-reranker-2b,
title={Supertron2-Reranker-2B -- Compact Cross-Encoder Reranking Model},
author={Surpem},
year={2026},
url={https://huggingface.co/Surpem/Supertron2-Reranker-2B},
}