Skip to content

tiny-research

QWEN-0.5B-GRPO

tiny-research/QWEN-0.5B-GRPO

QWEN-0.5B-GRPO is a machine learning model from tiny-research. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.

Finetuned Qwen2.5-0.5B on GSM8k, using Group Relative Policy Optimization proposed on DeepSeekMath.

Downloads · 30 days

0

Access

Public

Updated Feb 13, 2025

Repo size

8.9 GB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.pt5.9 GB · 66%

At a glance

License
mit
Access
Public
Created
Feb 13, 2025
Updated
Feb 13, 2025
SHA
3f8b3509
License
mit
Created
Feb 13, 2025
Updated
Feb 13, 2025