Downloads · 30 days
96
17% of all-time downloads
ForSureTesterSim/Qwen2.5-R1-Minny-1.5B-v1
Qwen2.5-R1-Minny-1.5B-v1 is a text generation model from ForSureTesterSim. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Qwen2.5-R1-Minny-1.5B-v1 is a highly plastic, multi-domain "Super-Base" model engineered through a novel hybrid merging topology. It fuses state-of-the-art latent reasoning (GRPO), mathematical logic, and coding profi…
Downloads · 30 days
96
17% of all-time downloads
All-time downloads
558
Public
Parameters
1.8B
3.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.6 GB · 100%
From the Hugging Face model README
Qwen2.5-R1-Minny-1.5B-v1 is a highly plastic, multi-domain "Super-Base" model engineered through a novel hybrid merging topology. It fuses state-of-the-art latent reasoning (GRPO), mathematical logic, and coding proficiency into a single 1.5B parameter manifold without suffering from catastrophic forgetting or RLHF-induced formatting collapse.
This model was created using a custom out-of-core pipeline combining Model Stock (geometric anchoring) and Sens-Merging (gradient-based sensitivity scaling), informed by extensive layer-wise ablation studies.
This model is the result of the "Refined Golden Triangle" merging topology, which anchors on a Group Relative Policy Optimization (GRPO) reasoning model and injects pure Supervised Fine-Tuned (SFT) domain experts.
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5BRLinf/RLinf-math-1.5Bagentica-org/DeepCoder-1.5B-Previewmobiuslabsgmbh/DeepSeek-R1-ReDistill-Qwen-1.5B-v1.1During the development of this model, extensive forward-pass perplexity benchmarking revealed several fundamental laws regarding the merging of modern RLHF models:
To bypass these limitations, Qwen2.5-R1-Minny-1.5B-v1 utilizes a custom hybrid algorithm:
Prior to minting this model, it was evaluated against raw baseline algorithms using a strict, zero-shot Forward-Pass Cross-Entropy Loss evaluation across three domains (lower loss = better absorption of the latent space):
| Merging Algorithm | Foundation Loss (FineWeb) ↓ | Reasoning Loss (OpenMath) ↓ | Code Loss (CodeAlpaca) ↓ |
|---|---|---|---|
| Raw Task Arithmetic | 3.8883 | 1.8376 | 2.2205 |
| Raw Model Stock | 3.7151 | 1.7052 | 1.9991 |
| Model Stock + Sens-Merge (This Model) | 3.5973 | 1.5694 | 1.6017 |
By combining the geometric stability of Model Stock with the gradient-aware routing of Sens-Merging, this model achieves the strict Pareto-optimal frontier for 1.5B reasoning architectures.
Qwen2.5-R1-Minny-1.5B-v1 is primarily intended to serve as a Super-Base for downstream alignment.
While it can be used for zero-shot inference, its true potential is unlocked when used as the initialization checkpoint for downstream Supervised Fine-Tuning (SFT) or Direct Preference Optimization (DPO). Because the Base Anchor's foundational neurons were mathematically shielded during the merge, the model retains maximum plasticity. It holds deep latent representations of Python/C++ coding syntax, mathematical routing, and DeepSeek <think> traces, ready to be aligned to your specific conversational formatting or agentic workflow.
<|im_start|> vs <|eot_id|>) may be slightly unstable out-of-the-box prior to downstream instruct-tuning.mergekit (Goddard et al.).