Downloads · 30 days
34
100% of all-time downloads
anonymous-ed-benchmark/SKILLRET-Reranker-0.6B
SKILLRET-Reranker-0.6B is a text ranking model from anonymous-ed-benchmark. Use it for the text ranking task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This is a reranker fine-tuned for AI agent skill retrieval. Given a natural-language user request and a candidate agent skill, it scores how relevant and useful the skill is for the request. It is designed as the seco…
Downloads · 30 days
34
100% of all-time downloads
All-time downloads
34
Public
Parameters
596M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 GB · 99%
From the Hugging Face model README
This is a reranker fine-tuned for AI agent skill retrieval. Given a natural-language user request and a candidate agent skill, it scores how relevant and useful the skill is for the request. It is designed as the second stage after a first-stage retriever such as SkillRet-Embedding-0.6B or SkillRet-Embedding-8B.
The model is fine-tuned from Qwen/Qwen3-Reranker-0.6B on the SkillRet benchmark training split with binary cross-entropy on the yes/no token probability. It keeps the scoring interface of Qwen3-Reranker.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "anonymous-ed-benchmark/SKILLRET-Reranker-0.6B"
tokenizer = AutoTokenizer.from_pretrained(model_id, padding_side="left")
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16).eval()
token_yes = tokenizer.convert_tokens_to_ids("yes")
token_no = tokenizer.convert_tokens_to_ids("no")
instruction = (
"Given a skill search query, judge whether the skill document "
"is relevant and useful for the query"
)
prefix = (
"<|im_start|>system\nJudge whether the Document meets the requirements based on the "
'Query and the Instruct provided. Note that the answer can only be "yes" or "no".'
"<|im_end|>\n<|im_start|>user\n"
)
suffix = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
prefix_ids = tokenizer.encode(prefix, add_special_tokens=False)
suffix_ids = tokenizer.encode(suffix, add_special_tokens=False)
max_length = 8192
def format_pair(query: str, doc: str) -> str:
return f"<Instruct>: {instruction}\n<Query>: {query}\n<Document>: {doc}"
@torch.no_grad()
def score(query: str, docs: list[str]) -> list[float]:
pairs = [format_pair(query, d) for d in docs]
enc = tokenizer(
pairs,
padding=False,
truncation="longest_first",
return_attention_mask=False,
max_length=max_length - len(prefix_ids) - len(suffix_ids),
)
enc["input_ids"] = [prefix_ids + ids + suffix_ids for ids in enc["input_ids"]]
enc = tokenizer.pad(enc, padding=True, return_tensors="pt").to(model.device)
logits = model(**enc).logits[:, -1, :]
stacked = torch.stack([logits[:, token_no], logits[:, token_yes]], dim=1)
return torch.nn.functional.log_softmax(stacked, dim=1)[:, 1].exp().tolist()
query = "Help me set up a CI/CD pipeline for my Python project"
skills = [
"ci-cd-setup | Configure continuous integration and deployment pipelines ...",
"python-debugging | Debug Python applications using pdb and logging ...",
]
print(score(query, skills)) # higher = more relevant
Each skill document is formatted as name | description | SKILL.md body, the same representation used by the SkillRet embedding models.
yes) for each query–skill pairEvaluated on the SkillRet benchmark evaluation split (4,392 queries, 6,006 skills). The reranker rescores the top-20 candidates returned by SkillRet-Embedding-8B.
| Model | NDCG@5 | NDCG@10 | NDCG@15 |
|---|---|---|---|
| SkillRet-Embedding-8B, no reranking | 0.8458 | 0.8644 | 0.8695 |
| + SkillRet-Reranker-0.6B (this model) | 0.8610 | 0.8774 | 0.8821 |
Full metrics for this model:
| Metric | @5 | @10 | @15 |
|---|---|---|---|
| NDCG | 0.8610 | 0.8774 | 0.8821 |
| Recall | 0.8928 | 0.9357 | 0.9510 |
| Completeness | 0.8206 | 0.8896 | 0.9128 |
This model is designed to rerank candidate agent skills for a natural-language user request. It is part of the SkillRet benchmark submission for evaluating skill retrieval systems for AI agents.
Citation information will be added in the de-anonymized release.