Downloads · 30 days
6.8K
11% of all-time downloads
sdadas/polish-reranker-roberta-v2
polish-reranker-roberta-v2 is a text ranking model from sdadas. Use it for the text ranking task on the model card, and read the license before you ship it in a product. It is set up for sentence-transformers. The card lists the license as gemma.
<h1 align="center"polish-reranker-roberta-v2</h1
Downloads · 30 days
6.8K
11% of all-time downloads
All-time downloads
61.2K
Public
Parameters
435M
870 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors870 MB · 98%
From the Hugging Face model README
This is an improved version of reranker based on sdadas/polish-roberta-large-v2 trained with RankNet loss on a large dataset of text pairs. The model was trained in the same way and on the same data as sdadas/polish-reranker-large-ranknet, but predictions from BAAI/bge-reranker-v2.5-gemma2-lightweight were used for distillation instead of unicamp-dl/mt5-13b-mmarco-100k.
Our reranker achieves results close to BAAI/bge-reranker-v2.5-gemma2-lightweight on the PIRB benchmark, even outperforming it on some datasets. At the same time, it is over 21 times smaller — 435M vs. 9.24B parameters.
The model can be used with Huggingface Transformers in the following way:
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import numpy as np
query = "Jak dożyć 100 lat?"
answers = [
"Trzeba zdrowo się odżywiać i uprawiać sport.",
"Trzeba pić alkohol, imprezować i jeździć szybkimi autami.",
"Gdy trwała kampania politycy zapewniali, że rozprawią się z zakazem niedzielnego handlu."
]
model_name = "sdadas/polish-reranker-roberta-v2"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(
model_name,
dtype=torch.bfloat16,
device_map="cuda"
)
texts = [f"{query}</s></s>{answer}" for answer in answers]
tokens = tokenizer(texts, padding="longest", max_length=512, truncation=True, return_tensors="pt").to("cuda")
output = model(**tokens)
results = output.logits.detach().cpu().float().numpy()
results = np.squeeze(results)
print(results.tolist())
The model achieves NDCG@10 of 65.30 in the Rerankers category of the Polish Information Retrieval Benchmark. See PIRB Leaderboard for detailed results.
@article{dadas2024assessing,
title={Assessing generalization capability of text ranking models in Polish},
author={Sławomir Dadas and Małgorzata Grębowiec},
year={2024},
eprint={2402.14318},
archivePrefix={arXiv},
primaryClass={cs.CL}
}