Downloads · 30 days
33
14% of all-time downloads
nbeerbower/Hemlock2-Coder-7B-GRPO
Hemlock2-Coder-7B-GRPO is a reinforcement learning model from nbeerbower. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
hemlang/Hemlock2-Coder-7B improved with execution-reward GRPO (grimoire ≥ 2.0 / hemlock-rl). Completions were executed in the Hemlock interpreter's sandbox and rewarded for exactly reproducing verified reference stdout.
Downloads · 30 days
33
14% of all-time downloads
All-time downloads
240
Public
Parameters
7.6B
23.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors15.2 GB · 65%
From the Hugging Face model README
hemlang/Hemlock2-Coder-7B improved with execution-reward GRPO (grimoire ≥ 2.0 / hemlock-rl). Completions were executed in the Hemlock interpreter's sandbox and rewarded for exactly reproducing verified reference stdout.
Strictly improves the base model on hembench (zero-shot, n=5, benchmark-overlapping training tasks held out):
| pass@1 | pass@5 | |
|---|---|---|
| Hemlock2-Coder-7B | 28.9% | 55.3% |
| Hemlock2-Coder-7B-GRPO | 36.8% (+7.9) | 57.9% (+2.6) |
Largest gains in syntax (L1 pass@1 2/9 → 4/9, pass@5 7/9 → 8/9) and systems/concurrency (L4 pass@1 1/7 → 3/7).
Q8_0 GGUF included.