Downloads · 30 days
14
0% of all-time downloads
lyndoncarlson/reranker
reranker is a machine learning model from lyndoncarlson. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This model is a reranker for Spanish text passages, built on top of BETO (a BERT-based model pre-trained on Spanish). It was trained to score the relevance of text passages given a user prompt, enabling you to reorder…
Downloads · 30 days
14
0% of all-time downloads
All-time downloads
30.5K
Public
Parameters
110M
439 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors439 MB · 100%
From the Hugging Face model README
This model is a reranker for Spanish text passages, built on top of BETO (a BERT-based model pre-trained on Spanish). It was trained to score the relevance of text passages given a user prompt, enabling you to reorder search results or candidate answers by how closely they match the user’s query.
reranker_beto_pytorch_optimized(prompt, content) pair, the model outputs a single numerical score indicating predicted relevance.Data Source:
(prompt, content, rank) tuples from the database.Preprocessing:
prompt, content) were tokenized using the BETO tokenizer (cased) with:
max_length = 512doc_stride = 256 (for lengthy passages)rank field was normalized and mapped to a continuous value (relevance) for regression.Training Setup:
dccuchile/bert-base-spanish-wwm-casedrelevance scoreAdamW with a learning rate of 3e-5Splits:
sklearn.model_selection.train_test_split.test_loss.Below is a quick example in Python using Hugging Face Transformers. After you’ve downloaded the model and tokenizer to ./reranker_beto_pytorch_optimized, you can do:
import torch
from transformers import BertTokenizer, BertForSequenceClassification
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
# Load the fine-tuned model and tokenizer
model_dir = "./reranker_beto_pytorch_optimized"
tokenizer = BertTokenizer.from_pretrained(model_dir)
model = BertForSequenceClassification.from_pretrained(model_dir).to(device)
model.eval()
prompt = "¿Cómo implementar un sistema solar en una escuela primaria?"
passage = "Este documento describe las partes del sistema solar ..."
inputs = tokenizer(
prompt,
passage,
max_length=512,
truncation='only_second',
padding='max_length',
return_tensors='pt'
)
# Forward pass
with torch.no_grad():
outputs = model(
input_ids=inputs['input_ids'].to(device),
attention_mask=inputs['attention_mask'].to(device)
)
score = outputs.logits.squeeze().item()
print(f"Predicted relevance score: {score:.4f}")
You would compare scores across multiple passages for a single prompt, then rank or sort them from highest to lowest predicted relevance.