Downloads · 30 days
11
7% of all-time downloads
dzhuj/reward
reward is a text classification model from dzhuj. Use it when you need a label for a piece of text. It is set up for transformers.
This model is a fine-tuned version of HuggingFaceTB/SmolLM-135M-Instruct on the HumanLLMs/Human-Like-DPO-Dataset dataset. It has been trained using TRL.
Downloads · 30 days
11
7% of all-time downloads
All-time downloads
151
Public
Parameters
135M
2.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors538 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of HuggingFaceTB/SmolLM-135M-Instruct on the HumanLLMs/Human-Like-DPO-Dataset dataset. It has been trained using TRL.
from transformers import pipeline
sample = [
{'content': 'Do you have a favorite type of music or artist?',
'role': 'user'},
{'content': "You know, I'm a big fan of indie-rock music. There's something about the raw, emotional vibe that really speaks to me. Arctic Monkeys are one of my all-time favorite bands - their lyrics are so clever and witty! But I'm also really into Tame Impala's psychedelic sound, it's like a trippy dream come true 😊. How about you, do you have a go-to genre or artist that gets you pumped up or relaxed? 🎵",
'role': 'assistant'}
]
reward_model = AutoModelForSequenceClassification.from_pretrained('dzhuj/reward')
tokenizer = AutoTokenizer.from_pretrained('dzhuj/reward')
inputs_sample = tokenizer.apply_chat_template(sample, tokenize=False)
inputs_sample = tokenizer(inputs_sample, return_tensors="pt")
score = reward_model(**inputs_sample).logits[0].cpu().detach().item()
This model was trained with Reward.