Downloads · 30 days
2
11% of all-time downloads
RTO-RL/Llama3-8B-RTO_RPP
Llama3-8B-RTO_RPP is a machine learning model from RTO-RL. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Base Model: OpenRLHF/Llama-3-8b-sft-mixture
Downloads · 30 days
2
11% of all-time downloads
All-time downloads
19
Public
Parameters
8B
16.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
Base Model: OpenRLHF/Llama-3-8b-sft-mixture
DPO model: RTO-RL/Llama3-8B-DPO
Reward model: RTO-RL/Llama3.2-1B-RewardModel
Prompt dataset: weqweasdas/ultra_train