Downloads · 30 days
0
ChrisMcCormick/ppo-tutorial-qwen2.5-0.5b
ppo-tutorial-qwen2.5-0.5b is a machine learning model from ChrisMcCormick. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Weights for a teaching notebook: Simple PPO — the advantage and loss for a single token (grpo-arithmetic course material). Qwen2.5-0.5B-Instruct repackaged as flat fp16 "parameter banks" (per-layer weights stacked alo…
Downloads · 30 days
0
Access
Public
Updated Aug 7, 2026
Repo size
4 GB
Likes
0
Public
Click a slice to open those files.
.safetensors3 GB · 100%
From the Hugging Face model README
Weights for a teaching notebook: Simple PPO — the advantage and loss for a
single token (grpo-arithmetic course material). Qwen2.5-0.5B-Instruct
repackaged as flat fp16 "parameter banks" (per-layer weights stacked along
dim 0) consumed by a ~60-line handwritten forward pass — no transformers
model classes involved.
| file | role in the notebook |
|---|---|
policy.safetensors | the model being trained (verl PPO checkpoint on GSM8K) |
reference.safetensors | frozen KL anchor: the original Qwen2.5-0.5B-Instruct |
value.safetensors | the critic: same trunk + scalar score head (verl AutoModelForTokenClassification layout) |
rollout.json | the fixed GSM8K question/response the notebook dissects |
fp16 (not bf16) so the notebook runs on a free-tier Colab T4.
Produced by agent-ops/grpo-arithmetic/2026-08-07_1043am_simple-ppo-notebook/
(convert_to_banks.py; parity + fp16-safety checks in parity_test.py).