Downloads · 30 days
0
kishan51/llm-zero-lite-experiments
llm-zero-lite-experiments is a reinforcement learning model from kishan51. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
A controlled comparison of continuous GRPO, fixed staged GRPO, and an LLM-controlled staged GRPO schedule on three-number Countdown using Qwen/Qwen3-1.7B with LoRA.
Downloads · 30 days
0
Access
Public
Updated Jun 21, 2026
Repo size
7.3 GB
Likes
0
Public
Click a slice to open those files.
.safetensors4 GB · 51%
From the Hugging Face model README
A controlled comparison of continuous GRPO, fixed staged GRPO, and an
LLM-controlled staged GRPO schedule on three-number Countdown using
Qwen/Qwen3-1.7B with LoRA.
| Method | Greedy accuracy | Sampled pass@1 | Sampled pass@4 |
|---|---|---|---|
| Continuous GRPO | 26.5% | 31.0% | 35.5% |
| Fixed staged GRPO | 34.5% | 34.5% | 39.5% |
| LLM controller | 36.5% | 37.5% | 40.5% |
The runs/ directory contains metrics, evaluation samples, configuration
history, controller decisions, logs, plots, and all saved LoRA checkpoints.