Downloads · 30 days
31
10% of all-time downloads
xinlai/DeepSeekMath-Base-SFT-Step-DPO
DeepSeekMath-Base-SFT-Step-DPO is a text generation model from xinlai. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
31
10% of all-time downloads
All-time downloads
305
Public
Parameters
6.9B
13.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors13.8 GB · 100%
From the Hugging Face model README
This repo contains the DeepSeekMath-Base-SFT-Step-DPO model. It is obtained by performing Step-DPO on DeepSeekMath-Base-SFT.
Step-DPO is a simple, effective, and data-efficient method for boosting the mathematical reasoning ability of LLMs. Notably, Step-DPO, when applied to Qwen2-72B-Instruct, achieves scores of 70.8% and 94.0% on the test sets of MATH and GSM8K without bells and wistles, respectively, surpassing a series of closed-source models, including GPT-4-1106, Claude-3-Opus, and Gemini-1.5-Pro.