Downloads · 30 days
3
9% of all-time downloads
explcre/scale-base-blend-dr-s50
scale-base-blend-dr-s50 is a reinforcement learning model from explcre. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
RL checkpoint. Base TikZilla-3B (SFT), trained with Dr.GRPO/DAPO against the DSV blenddr verifiable reward (0.7·DSV + 0.3·RenderGraph, both symbolic verifiers -- no learned reward model).
Downloads · 30 days
3
9% of all-time downloads
All-time downloads
33
Public
Parameters
3.1B
12.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6.2 GB · 100%
From the Hugging Face model README
RL checkpoint. Base TikZilla-3B (SFT), trained with Dr.GRPO/DAPO
against the DSV blend_dr verifiable reward (0.7·DSV + 0.3·RenderGraph, both symbolic verifiers --
no learned reward model).
explcre/datikz_v4_rl_hybrid_428k
(427,553 rows, 46.6% Kimi-K3 descriptions / 53.4% DaTikZ-v4 vlm_description, tagged per row)explcre/DeTikZify branch dsv-verifier @ c348406Evaluate with the matching input distribution. The same checkpoint scores differently on
vlmdesc984 vs kimidesc984, and differently again under different TeX/compile-gate versions --
always report the measurement config with the number.