Downloads · 30 days
0
gitcommit90/Qwen3.8-Flash-Next-NVFP4-DenseFP8-One-Spark
Qwen3.8-Flash-Next-NVFP4-DenseFP8-One-Spark is a image-text-to-text model from gitcommit90. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for vllm. The card lists the license as other.
A tested one-DGX-Spark deployment of nvidia/Qwen3.8-Flash-Next-NVFP4 with targeted in-loader FP8 conversion of the remaining large BF16 dense projections, native MTP K=3, piecewise CUDA graphs, prefix caching, and C6…
Downloads · 30 days
0
Access
Public
Updated Sep 6, 2026
Repo size
—
Likes
1
Trending 1
Click a slice to open those files.
.md4.9 KB · 76%
From the Hugging Face model README
A tested one-DGX-Spark deployment of nvidia/Qwen3.8-Flash-Next-NVFP4 with targeted in-loader FP8 conversion of the remaining large BF16 dense projections, native MTP K=3, piecewise CUDA graphs, prefix caching, and C6 scheduling.
47–53 tok/s short code · 37–43 tok/s long-context/thinking · 60 tok/s structured · 969K-token BF16 KV pool · 72/72 BENCHY low-effort cases across C1–C6
This page describes a runtime recipe; it does not yet contain transformed model weights. The dense projections are converted to FP8 during model loading. A stock-vLLM safetensors export is a separate planned artifact and is not claimed here.
The reproducible runtime, overlays, benchmark evidence, and setup scripts live in the companion GitHub repository: https://github.com/gitcommit90/qwen38-flash-next-nvfp4-one-spark.
NVIDIA's checkpoint stores the main routed experts in NVFP4 but leaves attention, Gated DeltaNet, lm_head, and other dense layers in BF16. Profiling on GB10 showed approximately 67% of decode kernel time in BF16 dense GEMVs.
The runtime recipe converts only the large decode-critical projections:
linear_attn.in_proj_qkvzlinear_attn.out_projself_attn.qkv_projself_attn.o_projlm_headSmall hyper-connection, shared-expert, gate, indexer, and auxiliary matrices remain BF16 because FP8 kernel overhead made them slower on this device. The selected kernel is Marlin FP8 W8A16: FP8 weights with BF16 activations.
Measured on one NVIDIA DGX Spark / GB10, TP1, 262,144-token model context, BF16 KV, MTP K=3, C6 scheduler capacity:
| Workload | Result |
|---|---|
| Short-code decode | 47–53 tok/s |
| Decode at 15K context | 37.2–37.8 tok/s |
| Thinking-on code | 39–43 tok/s |
| Open prose | 33.7 tok/s |
| Structured output | 60.2 tok/s |
| Cold prefill, 8–64K | ~1,980–2,082 tok/s |
| Warm cached 8K TTFT | 0.185–0.189 s |
| KV capacity | 969,224 tokens / ~3.70 × 262K |
The observed K3 agentic loop reached 74–81 decode tok/s, but shorter generated tool calls favored that figure. The robust comparison is 3.95 emitted tokens per target-model step at K3 versus 2.96 at K2 (+33%).
A 24-run C1–C6 × effort matrix was executed without retries, edits, or overwritten passes: 288 total cases.
| Reasoning effort | Outcome across C1–C6 |
|---|---|
| Low | 72 correct / 0 wrong / 0 unfinished |
| Medium | 65 correct / 0 wrong / 7 unfinished |
| XHigh | 60 correct / 3 wrong / 9 unfinished |
| None | 13 correct / 59 wrong / 0 unfinished |
Recommended default for agent workloads: reasoning_effort: low.
Teacher-forced target-logit comparison against the NVIDIA checkpoint with BF16 dense layers, 30 texts / 2,019 positions:
No fine-tuning was performed. Native MTP speculation verifies candidate tokens against the transformed target.
fab0aecb760cec45227f6656abcaafa11abca87a8a728663c1c3eeace834a95f5654fa653cc1998ce648f1a131365aae15920073e761a3fa5a527654gdn,attn,lm_headLocal text, coding, tool-use, and agent workloads on one DGX Spark where decode latency, usable concurrency, and large BF16 KV capacity are more important than retaining byte-identical BF16 dense weights.
The runtime code is Apache-2.0. Upstream model weights remain governed by the NVIDIA Open Model License and applicable Qwen terms. This repository does not redistribute those weights.