Downloads · 30 days
75
1% of all-time downloads
cloudyu/19B_MATH_DPO
19B_MATH_DPO is a text generation model from cloudyu. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
this is a DPO fine-tuned MoE model with about 19B parameter.
Downloads · 30 days
75
1% of all-time downloads
All-time downloads
10.6K
Public
Parameters
19.2B
38.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors38.4 GB · 100%
From the Hugging Face model README
this is a DPO fine-tuned MoE model with about 19B parameter.
DPO Trainer
TRL supports the DPO Trainer for training language models from preference data, as described in the paper Direct Preference Optimization: Your Language Model is Secretly a Reward Model by Rafailov et al., 2023.