Downloads · 30 days
79
26% of all-time downloads
patrickbdevaney/DeepSeek-V4-Flash-0731-REAP-DSpark
DeepSeek-V4-Flash-0731-REAP-DSpark is a text generation model from patrickbdevaney. Use it when you need the model to write or continue text. It is set up for safetensors. The card lists the license as mit.
A drop-in replacement for 0xSero/DeepSeek-V4-Flash-0731-REAP with a fine-tuned DSpark MTP draft head merged in. Download it and you have the whole model: no patching step, no separate head file to graft on.
Downloads · 30 days
79
26% of all-time downloads
All-time downloads
302
Public
Parameters
193B
108 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors108 GB · 100%
How the weights are stored.
I8185B · 96%
From the Hugging Face model README
A drop-in replacement for 0xSero/DeepSeek-V4-Flash-0731-REAP
with a fine-tuned DSpark MTP draft head merged in. Download it and you have the whole model: no
patching step, no separate head file to graft on.
Every emitted token is bit-identical to base autoregressive decode. The speed-up is lossless by construction — speculation changes how tokens are produced, never which tokens they are — and that invariant is checked on every run by a gate that compares the speculative path against plain AR.
Three files. Shards 46, 47 and 48 carry the MTP draft-head tensors; shards 1–45, the tokenizer, and the config are byte-identical to the base. If you already have the base checkpoint, the only new weights here are ~6.7 GB of head shards.
| provenance | deepseek-ai/DeepSeek-V4-Flash-0731 @ 9e165c30 → REAP prune (256 → 160 experts/layer) → this head fine-tune |
| head trained on | 1,472 self-generated sequences, deficit-weighted, a_ce 0.1 / a_tv 0.9 |
| draft block width | 5 (the checkpoint's own dspark_block_size) |
| quantisation | native MXFP4 experts, FP8 attention — unchanged from base |
Measured on a Jetson AGX Thor (sm_110a, 122.8 GiB unified) with a from-scratch pure-CUDA
inference server. Every figure is from a logged run, not an estimate.
| measured | |
|---|---|
suite mean tau | 3.8413 of a 5-wide draft (77 % of the ceiling) |
| suite mean throughput | 28.38 tok/s |
| base AR decode | 14.61 tok/s → speculation is 1.94× |
| vs the stock shipped head | 22.66 → 28.38 tok/s, +25.3 % |
| on copy-heavy agentic prompts | up to 34.97 tok/s, 2.51× |
tau is tokens committed per target forward. The head was selected on the mean, and the mean
hides a real trade — published here rather than buried:
| category | stock head | this head | |
|---|---|---|---|
| long_context | 5.00 | 4.53 | −0.47 |
| agentic_format | 4.41 | 4.16 | −0.25 |
| multi_turn | 4.06 | 4.83 | +0.77 |
| code_edit | 4.08 | 4.18 | +0.10 |
| short_factual | 3.11 | 3.87 | +0.76 |
| reasoning | 1.82 | 3.57 | +1.75 |
| explanation | 1.70 | 3.00 | +1.30 |
| code_gen | 1.86 | 2.59 | +0.73 |
| mean | 3.2550 | 3.8413 | +18.0 % |
The gain is concentrated in the constructive categories — reasoning nearly doubles — and it is paid for on the two reconstructive ones the stock head was already best at. Under this project's own release rule that is a failure of 3 of 6 category floors, and the rule is stated so that a downstream user can predict their own workload instead of inheriting a mean.
If your workload is dominated by long-context reconstruction, the stock head may serve you better. For mixed agentic work — tool calls, multi-turn, code edits, reasoning — this head is a clear win.
Measured through the same server, effort=low, 24k token budget, temperature 1.0 / top-p 0.95.
| benchmark | score | n |
|---|---|---|
| AIME 2024 | 91.7 % [80.0, 100.0] | 60 |
| GPQA-Diamond | 83.8 % [78.1, 88.3] | 198 |
| MMLU-Pro | 74.7 % [67.2, 81.0] | 150 |
| BFCL (tool calling) | 86.2 % | 240 |
| BFCL-Live | 78.7 % | 508 |
| LiveCodeBench | 46.9 % [39.6, 54.2] (8k budget, 59 % truncated — not quotable) | 175 |
A caveat that matters more than any single number. GPQA at a 24k budget was pre-registered at 87–93 % before the run, with a stated falsification threshold of 85 %. It landed at 83.8 % — the prediction was wrong, and the reason is a finding about this pruned model: 19 of 51 items that exhaust an 8k budget never terminate even at 24k, and the ones that do finish score 78.1 %, not the 93.9 % that items terminating inside 8k reach. Prompts needing long chains are qualitatively harder here, not merely longer. Budget your context accordingly.
| context | KV format | fits |
|---|---|---|
| 32,768 | fp32 rows | yes, 119.9 / 122.8 GiB |
| 131,072 | packed fp8 + ue8m0 | yes, 117.0 / 122.8 GiB |
| 262,144 | packed | allocates at 121.8 / 122.8 — too little headroom to serve |
KV packing is bit-exact (720 B/row vs 2048; the fp32 path already stored e4m3-representable values in four bytes) and costs ~12 % of prefill throughput, so it is worth enabling only when the context needs it.
MIT, following the base checkpoint. The DeepSeek copyright notice is retained in LICENSE. This
repository adds only the fine-tuned draft-head tensors in shards 46–48.