Downloads · 30 days
13
48% of all-time downloads
RTO-RL/Llama3-8B-RewardModel
Llama3-8B-RewardModel is a machine learning model from RTO-RL. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Base model: OpenRLHF/Llama-3-8b-sft-mixture
Downloads · 30 days
13
48% of all-time downloads
All-time downloads
27
Public
Parameters
7.5B
15 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors15 GB · 100%
From the Hugging Face model README
Base model: OpenRLHF/Llama-3-8b-sft-mixture
Preference dataset: HuggingFaceH4/ultrafeedback_binarized