Downloads · 30 days
0
tiny-research/QWEN-0.5B-GRPO
QWEN-0.5B-GRPO is a machine learning model from tiny-research. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Finetuned Qwen2.5-0.5B on GSM8k, using Group Relative Policy Optimization proposed on DeepSeekMath.
Downloads · 30 days
0
Access
Public
Updated Feb 13, 2025
Repo size
8.9 GB
Likes
0
Public
Click a slice to open those files.
.pt5.9 GB · 66%
From the Hugging Face model README
Finetuned Qwen2.5-0.5B on GSM8k, using Group Relative Policy Optimization proposed on DeepSeekMath.
A random sample,
Question:
Jenna is adding black dots to a bunch of white blouses. Each blouse gets 20 dots, and each dot takes 10 ml of black dye. How many 400-ml bottles of dye does Jenna need to buy to dye 100 blouses?
Answer:
50
Response:
<reasoning>
Jenna needs to dye 100 blouses, and each blouse gets 20 dots, so she needs 100 * 20 = 2000 dots in total. Each dot takes 10 ml of black dye, so she needs 2000 * 10 = 20000 ml of black dye. Since each bottle holds 400 ml of dye, Jenna needs to buy 20000 / 400 = 50 bottles of dye.
</reasoning>
<answer>
50
</answer>