Downloads · 30 days
0
Aravind0495/AetherControl-Qwen2.5-1.5B-GRPO-Math
AetherControl-Qwen2.5-1.5B-GRPO-Math is a machine learning model from Aravind0495. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This model is fine-tuned using Group Relative Policy Optimization (GRPO) on 500 reasoning prompts from the GSM8K dataset as part of the AetherControl platform.
Downloads · 30 days
0
Access
Public
Updated Jul 25, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md2.6 KB · 63%
From the Hugging Face model README
This model is fine-tuned using Group Relative Policy Optimization (GRPO) on 500 reasoning prompts from the GSM8K dataset as part of the AetherControl platform.
📊 GRPO Fine-Tuning Progression & Telemetry Log
┏━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Step ┃ Mean Reward ┃ Format ┃ Accuracy ┃ KL ┃ GRPO Loss ┃
┃ ┃ (r_mean) ┃ Reward ┃ Reward ┃ Div (D_KL) ┃ (L_grpo) ┃
┡━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Step 01 │ 0.42 │ 0.23 │ 0.22 │ 0.0376 │ 1.1934 │
│ Step 05 │ 0.50 │ 0.34 │ 0.32 │ 0.0581 │ 0.9610 │
│ Step 10 │ 0.64 │ 0.54 │ 0.48 │ 0.0595 │ 0.7322 │
│ Step 15 │ 0.72 │ 0.75 │ 0.70 │ 0.0536 │ 0.4367 │
│ Step 20 │ 0.85 │ 0.92 │ 0.87 │ 0.0586 │ 0.1730 │
└─────────┴─────────────┴─────────────┴─────────────┴─────────────┴────────────┘
<think> tags) and exact math string match (math_reward.py).Github Repository: https://github.com/Aravind0403/Building-and-Optimizing-Production-LLM-Serving-System