Downloads · 30 days
1.5K
2% of all-time downloads
danielchalef/Qwen3-Reranker-4B-seq-cls-vllm-fixed
Qwen3-Reranker-4B-seq-cls-vllm-fixed is a text classification model from danielchalef. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
This is a fixed version of the Qwen3-Reranker-4B model converted to sequence classification format, optimized for use with vLLM.
Downloads · 30 days
1.5K
2% of all-time downloads
All-time downloads
73.3K
Public
Parameters
4B
16.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
This is a fixed version of the Qwen3-Reranker-4B model converted to sequence classification format, optimized for use with vLLM.
This model is a pre-converted version of Qwen/Qwen3-Reranker-4B that:
The original converted model (tomaarsen/Qwen3-Reranker-4B-seq-cls) was missing critical vLLM configuration attributes. This version adds:
{
"classifier_from_token": ["no", "yes"],
"method": "from_2_way_softmax",
"use_pad_token": false,
"is_original_qwen3_reranker": false
}
These configurations are essential for vLLM to properly handle the pre-converted weights.
vllm serve danielchalef/Qwen3-Reranker-4B-seq-cls-vllm-fixed \
--task score \
--served-model-name qwen3-reranker-4b \
--disable-log-requests
from vllm import LLM
llm = LLM(
model="danielchalef/Qwen3-Reranker-4B-seq-cls-vllm-fixed",
task="score"
)
queries = ["What is the capital of France?"]
documents = ["Paris is the capital of France."]
outputs = llm.score(queries, documents)
scores = [output.outputs.score for output in outputs]
print(scores)
This model performs identically to the original Qwen3-Reranker-4B when used with proper configuration, while providing significant efficiency improvements:
If you use this model, please cite the original Qwen3-Reranker:
@misc{qwen3reranker2024,
title={Qwen3-Reranker},
author={Qwen Team},
year={2024},
publisher={Hugging Face}
}
Apache 2.0 (inherited from the base model)