Downloads · 30 days
194
100% of all-time downloads
R3n3r0/dapack-math
dapack-math is a text generation model from R3n3r0. Use it when you need the model to write or continue text. The card lists the license as mit.
Qwen3.5-35B-A3B compressed for mathematics and quantitative reasoning: 10.7 GB instead of 21.2 GB, with every expert still present. The 143 experts this domain routes to are kept at q2K; the remaining 113 are held at…
Downloads · 30 days
194
100% of all-time downloads
All-time downloads
194
Public
Repo size
11.3 GB
Likes
0
Public
Click a slice to open those files.
.gguf11.3 GB · 100%
From the Hugging Face model README
Qwen3.5-35B-A3B compressed for mathematics and quantitative reasoning: 10.7 GB instead of 21.2 GB, with every expert still present. The 143 experts this domain routes to are kept at q2_K; the remaining 113 are held at IQ2_XXS (2.06 bpw) under an importance matrix — graded, not deleted, so the pack degrades out of domain instead of breaking.
⚠️ This file requires the dapack runtime
Graded packs carry two expert tensors per layer at different precisions (
ffn_*_exps_cold+dapack.hot_experts_per_layer). Stock llama.cpp / ollama / LM Studio cannot load them. Build the runtime from:https://github.com/R3n3r0/dapack — ready-to-run binaries under Releases (Linux x86_64, ROCm), nothing to compile:
tar xzf dapack-v0.1.0-linux-x86_64-rocm-gfx1151.tar.gz cd dapack-v0.1.0-* && ./dapack serve ~/catalog.dapack # web chat + OpenAI API + routingOn other GPUs, build from the same repo (one command,
scripts/rebuild_fork.sh --applyafter a recursive clone).
These numbers are behavioural measurements, not estimates. They ship in the pack's manifest, and the dapack router treats them as hard constraints — a request needing a capability this pack has lost is routed to a pack that has it.
| capability | this pack | full model (21.2 GB) |
|---|---|---|
| reasoning (GSM8K convergence, 4096-token budget, n=100) | 89.0% | 77.0% |
| tool calling (8 probes) | 8/8 | 8/8 |
| code generation | 100% | 100% |
| instruction following | 100% | 80% |
| long context (needle @ 3k tokens) | 100% | 100% |
| structured output (JSON) | 100% | 100% |
| translation en→it | 20% ⚠️ | 100% |
| Italian | not verified ⚠️ | verified |
What it lost — on purpose, and declared: translation collapsed (100% → 20%) and the
pack failed Italian certification outright, because a mathematics calibration held the
language experts at 2 bits. Its manifest says languages: [en], and under the dapack
router non-English requests never reach it. If you serve this file standalone, use it in
English.
On the identical expert selection, we measured:
| mechanism | Qwen3.5-35B | Qwen3-30B | tools |
|---|---|---|---|
| surplus experts deleted | 79.0% | 40.0% | 7/8 |
| surplus experts at 2 bits | 82.0% | 93.3% | 8/8 |
On the second architecture deletion loses 53 points and grading loses none. (35B columns measured on this pack's language-domain twin; same base, same recipe.)
Deletion costs ~1 point of reasoning per 1% of experts chosen wrongly and is unrecoverable. A 2-bit expert is present and merely imprecise. Compute cost is unchanged: the top-k budget is partitioned across the two banks, so exactly 8 experts run per token.
| file | size | needs |
|---|---|---|
graded_math.gguf | 10.7 GB | dapack fork |
manifest-fragment.json | — | capability manifest for the catalogue |
Full documentation, tools to build your own domain packs, and every measurement behind this card: https://github.com/R3n3r0/dapack