Downloads · 30 days
997
100% of all-time downloads
canada-quant/GLM-5.3-Flash-DFlash2-F
GLM-5.3-Flash-DFlash2-F is a text generation model from canada-quant. Use it when you need the model to write or continue text. It is set up for vllm. The card lists the license as apache-2.0.
Self-trained DFlash2 block-diffusion speculative-decoding drafter for GLM-5.3-Flash, trained against our INT4 quant canada-quant/GLM-5.3-Flash-W4A16-MTP. Trained on self-generated data only — no third-party drafter we…
Downloads · 30 days
997
100% of all-time downloads
All-time downloads
997
Public
Parameters
3.1B
6.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6.2 GB · 100%
From the Hugging Face model README
dflash2FSelf-trained DFlash2 block-diffusion speculative-decoding drafter for GLM-5.3-Flash,
trained against our INT4 quant canada-quant/GLM-5.3-Flash-W4A16-MTP.
Trained on self-generated data only — no third-party drafter weights or traces anywhere in the training path. Supersedes
GLM-5.3-Flash-DFlash2-E (same architecture, same serve contract, drop-in):
+0.065 mean acceptance at K=7 on the same hardware, and parity with the reference drafter
incoai/GLM-5.3-Flash-DFlash2 measured on the same GPUs the same day (see below).
-E)| Field | Value |
|---|---|
| Type | DFlash2 block-diffusion drafter (DFlash2DraftModel, Qwen3-style backbone) |
| Layers | 8 decoder layers, full attention (no sliding window) |
| Hidden / heads | 4096 (intermediate 12288) · 32 attention heads / 8 KV heads · head_dim 128 |
| Target taps | 9 hidden-state taps at target layers [5, 9, 14, 19, 24, 28, 33, 38, 42] |
| Block size | 8 → K = 7 speculative tokens (num_speculative_tokens: 7) |
| Selector | rank 256, top_k 16, grouped dynamic conv (kernel 2, group 16) — trained, not vestigial |
| Mask embedding | learnable, shipped as mask_embedding.pt (mask_token_id 154856) |
| Size | ~1.84B drafter parameters · 6.2 GB bf16 checkpoint (ships untied embed_tokens + lm_head) |
| Max positions | 1,048,576 |
dflash2E; one pass over 449,600 self-generated samples (the 350,260 of -E + 99,340 new completions of
never-before-generated prompts), 41,279 steps, lr 1e-4 → 1e-5 (cosine), gamma-6.5 tail weighting, CE + top-20 KL, 8× NVIDIA B300, 16.5 h.PROVENANCE.txt is the verbatim training record. Recipe, code, patches and every raw measurement:
canada-quant/vllm-glm53-flash-sm121/drafter (eval kit, every raw measurement, cards).500 never-trained-on prompts, thinking ON, T=1.0 / top_p 0.95, max_tokens 1024, greedy drafts, TP=4 (c16; the c1 row is 100 prompts at
concurrency 1). All rows below were measured on the same 8× B300 within six hours of each other, with the same vLLM build and the same
target backend (FLASHINFER_MLA_SPARSE) — acceptance length shifts by a few hundredths between GPU generations (incoai reads 3.602–3.615
on H200 vs 3.632 on B300), so only same-hardware rows are compared.
| Mean acceptance length (output tok/s) | K=7, c16 | K=4, c16 | K=7, c1 |
|---|---|---|---|
dflash2F (this) | 3.626 (1,406) | 3.085 (1,339) | 3.677 (313) |
| incoai/GLM-5.3-Flash-DFlash2 (CC-BY-NC-ND-4.0) | 3.632 (1,425) | 3.136 (1,369) | 3.667 (328) |
dflash2E (our previous release) | 3.561 (1,397) | — | — |
Per-position acceptance (K=7, c16): 0.750 · 0.555 · 0.415 · 0.316 · 0.244 · 0.193 · 0.153 (incoai: 0.75 · 0.55 · 0.41 · 0.32 · 0.25 · 0.20 · 0.16).
Honest read: parity with the reference at K=7 (−0.006) and single-stream (+0.010), both inside the protocol's ±0.015 run-to-run noise;
1.6% short at K=4 (−0.051, the first draft positions). Against -E it is +0.065 at K=7 — the largest single-leg gain in the lineage, from
100K fresh self-generated samples. We ship it because it is fully ours (Apache-2.0, no NC/ND terms), trained on data we control,
reproducible end-to-end, and now level with the reference. On the H200 (previous campaign) -E measured 3.568 / 3.067 / 3.585.
DGX Spark (SM121): -E reproduced its H200 acceptance on 2× Spark within 0.3% (3.5788 c16 / 3.5970 c1 vs 3.568 / 3.585). -F is a
drop-in for the same launcher; its Spark same-rig A/B against -E and the reference is the next measurement — re-measure both drafters on
your rig with the eval kit in the repo before quoting a Spark number.
Pairs with canada-quant/GLM-5.3-Flash-W4A16-MTP (the quant it was trained against) or BF16 GLM-5.3-Flash. vLLM speculative config:
{"method": "dflash", "model": "/models/GLM-5.3-Flash-DFlash2-F", "num_speculative_tokens": 7}
x86 (H100 / H200 / B300), upstream vLLM nightly ≥ 2026-09-08 (+ the two GLM-5.3-Flash DFlash2 bind-mount patches in the repo above):
vllm serve /models/glm53-flash-w4a16-mtp --served-model-name glm53 --tensor-parallel-size 4 --enable-expert-parallel \
--block-size 64 --no-enable-prefix-caching \
--speculative-config '{"method":"dflash","model":"/models/GLM-5.3-Flash-DFlash2-F","num_speculative_tokens":7}'
2× DGX Spark (SM121): prebuilt image + one-command launcher in
canada-quant/vllm-glm53-flash-sm121 — the drafter is bind-mounted at runtime
(DRAFTER_HOST_PATH=/models/GLM-5.3-Flash-DFlash2-F), no rebuild. launch_dflash2_tp2.sh in this repo is the
same TP=2 launcher with this drafter as its default.
Hard constraints (all measured, not stylistic):
num_speculative_tokens: 7 (= block_size − 1) is the trained and measured value. k=5 also boots on the SM121 image (2026-09-26, not
benchmarked); an earlier version of this card said other counts boot-wedge the stack, which was wrong.mask_embedding.pt must sit next to the weights. Verify the boot log carries Loaded DFlash mask embedding for mask_token_id 154856 from mask_embedding.pt — absence means the mask was silently ignored; do not serve.dflash_config.target_layer_ids of length 9 (upstream vLLM DFlash2 does —
vllm-project/vllm#52816).FullAttentionSpec drafter KV in the GLM-5 KV fast path (the SM121 image above
has it; upstream nightly needs the kv_cache_utils patch from the repo).This drafter's 8 layers use full attention, so each keeps a KV cache for the whole context (same architecture as -E / -G); the incoai reference drafter keeps a
2,048-token sliding window. KV pools measured by the authors on 2× DGX Spark (fp8 KV):
| drafter | KV pin | max model len | KV pool | date |
|---|---|---|---|---|
| incoai reference | 9 GiB | 1,048,576 | 1,360,420 tokens (≈7.1 KB/token) | 2026-08-31 |
-E / -G | 8 GiB | 262,144 | 366,749 tokens (≈22 KB/token) | 2026-09-16 |
-G | 16 GiB | 800,000 | 888,729 tokens (≈19 KB/token) | 2026-09-26 |
That is about 3× less context per GiB of KV. A 9 GiB pin holds ≈450K tokens with -E / -F / -G — less than one 1M request — so
1M context on 2× Spark is validated only with the incoai drafter; with our drafters the validated contexts are 262K (8 GiB) and 800K
(16 GiB).
{"method":"mtp","num_speculative_tokens":3}) for long
context, non-English text, or KV headroom.-G, 2× Spark, our SM121 image, single stream, llama-benchy pp2048/tg128): 31.4 / 30.7 / 15.3 / 10.0
tok/s at context depth 0 / 4K / 65K / 100K (2026-09-28; the 100K cell from 2026-09-26).-G usable context fell
from ≈930K to ≈330K tokens and decode was 20–30% slower on non-English text (faster on English code).canada-quant/vllm-glm53-flash-sm121.model.safetensors (sha256 305efe2aa50447cbac7d7d3ab566985103973a79c5d21eb11e1dbf870139dd8c), config.json (699bdf29…c51778, identical
to -E), mask_embedding.pt (0a5eecd9…304f3a), PROVENANCE.txt (546442a6…f75721).angelspec-glm53/runs/all/dflash2F-*.json,
dflash2E-k7-b300.json, incoai-k7-b300.json, incoai-k4-b300.json, incoai-k7-c1-b300.json).