Skip to content

D-URing

Hint_Model_GRPO_V1

D-URing/Hint_Model_GRPO_V1

Hint_Model_GRPO_V1 is a machine learning model from D-URing. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.

GRPO-trained Qwen3-4B-Instruct-2507 with HARPM (Hard-problem Adaptive Reference-Prompt Matching) hint injection.

Downloads · 30 days

5

17% of all-time downloads

All-time downloads

30

Public

Parameters

4B

8.1 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors8 GB · 100%

At a glance

License
other
Model type
qwen3
Access
Public
Created
Jul 13, 2026
Updated
Jul 13, 2026
SHA
716a9c72

Base models

Type
qwen3
License
other
Created
Jul 13, 2026
Updated
Jul 13, 2026
Hint_Model_GRPO_V1 — AI Model — AIMarketly