Skip to content

explcre

scale-base-blend-dr-s50

explcre/scale-base-blend-dr-s50

scale-base-blend-dr-s50 is a reinforcement learning model from explcre. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.

RL checkpoint. Base TikZilla-3B (SFT), trained with Dr.GRPO/DAPO against the DSV blenddr verifiable reward (0.7·DSV + 0.3·RenderGraph, both symbolic verifiers -- no learned reward model).

Downloads · 30 days

3

9% of all-time downloads

All-time downloads

33

Public

Parameters

3.1B

12.4 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors6.2 GB · 100%

At a glance

Task
Reinforcement Learning
License
mit
Model type
qwen2
Access
Public
Created
Jul 28, 2026
Updated
Aug 5, 2026
SHA
f2f2cc93
Task
Reinforcement Learning
Type
qwen2
License
mit
Created
Jul 28, 2026
Updated
Aug 5, 2026
scale-base-blend-dr-s50 — AI Model — AIMarketly