Downloads · 30 days
76
0% of all-time downloads
yunconglong/7Bx4_DPO
7Bx4_DPO is a text generation model from yunconglong. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
DPO Trainer with jondurbin/truthy-dpo-v0.1
Downloads · 30 days
76
0% of all-time downloads
All-time downloads
32.8K
Public
Parameters
24.2B
48.3 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors48.3 GB · 100%
From the Hugging Face model README
DPO Trainer TRL supports the DPO Trainer for training language models from preference data, as described in the paper Direct Preference Optimization: Your Language Model is Secretly a Reward Model by Rafailov et al., 2023.
"num_experts_per_tok": 4