Skip to content

nbeerbower

Hemlock2-Coder-7B-GRPO

nbeerbower/Hemlock2-Coder-7B-GRPO

Hemlock2-Coder-7B-GRPO is a reinforcement learning model from nbeerbower. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.

hemlang/Hemlock2-Coder-7B improved with execution-reward GRPO (grimoire ≥ 2.0 / hemlock-rl). Completions were executed in the Hemlock interpreter's sandbox and rewarded for exactly reproducing verified reference stdout.

Downloads · 30 days

33

14% of all-time downloads

All-time downloads

240

Public

Parameters

7.6B

23.3 GB on disk

Likes

1

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors15.2 GB · 65%

At a glance

Task
Reinforcement Learning
License
apache-2.0
Model type
qwen2
Access
Public
Created
Jul 15, 2026
Updated
Jul 15, 2026
SHA
13808029

Base models

Task
Reinforcement Learning
Type
qwen2
License
apache-2.0
Created
Jul 15, 2026
Updated
Jul 15, 2026
Hemlock2-Coder-7B-GRPO — AI Model — AIMarketly