Downloads · 30 days
5
17% of all-time downloads
D-URing/Hint_Model_GRPO_V1
Hint_Model_GRPO_V1 is a machine learning model from D-URing. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
GRPO-trained Qwen3-4B-Instruct-2507 with HARPM (Hard-problem Adaptive Reference-Prompt Matching) hint injection.
Downloads · 30 days
5
17% of all-time downloads
All-time downloads
30
Public
Parameters
4B
8.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
GRPO-trained Qwen3-4B-Instruct-2507 with HARPM (Hard-problem Adaptive Reference-Prompt Matching) hint injection.
| metric | untrained | baseline02 (15ep plain) | HINT (this model) |
|---|---|---|---|
| acc mean@4 | 0.026 | 0.067 | 0.0865 |
| acc best@4 | 0.046 | — | 0.122 |
Equal-budget improvement of +29% mean@4 over the plain baseline; validation accuracy increased monotonically over training.
Note: single-seed run; multi-seed variance not yet measured.