Downloads · 30 days
76
1% of all-time downloads
yunconglong/13B_MATH_DPO
13B_MATH_DPO is a text generation model from yunconglong. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
DPO Trainer with dataset kyujinpy/orcamathdpo to improve [yunconglong/MoE13BDPO]
Downloads · 30 days
76
1% of all-time downloads
All-time downloads
11.7K
Public
Parameters
12.9B
25.8 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors25.8 GB · 100%
From the Hugging Face model README
DPO Trainer TRL supports the DPO Trainer for training language models from preference data, as described in the paper Direct Preference Optimization: Your Language Model is Secretly a Reward Model by Rafailov et al., 2023.