Downloads · 30 days
4
0% of all-time downloads
SpiceRL/DRA-GRPO
DRA-GRPO is a machine learning model from SpiceRL. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
This model is described in the paper DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models.
Downloads · 30 days
4
0% of all-time downloads
All-time downloads
852
Public
Parameters
1.8B
3.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors3.6 GB · 100%
From the Hugging Face model README
This model is described in the paper DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models.
Full code is in: https://github.com/xiwenc1/DRA-GRPO