Downloads · 30 days
44
1% of all-time downloads
RTO-RL/Llama3.2-1B-RewardModel
Llama3.2-1B-RewardModel is a machine learning model from RTO-RL. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Base model: unsloth/Llama-3.2-1B-Instruct
Downloads · 30 days
44
1% of all-time downloads
All-time downloads
5.6K
Public
Parameters
1.2B
2.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.5 GB · 99%
From the Hugging Face model README
Base model: unsloth/Llama-3.2-1B-Instruct
Tokenizer: OpenRLHF/Llama-3-8b-sft-mixture
Preference dataset: HuggingFaceH4/ultrafeedback_binarized