Downloads · 30 days
149
6% of all-time downloads
modal-labs/Inkling-NVFP4-DFlash
Inkling-NVFP4-DFlash is a text generation model from modal-labs. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
149
6% of all-time downloads
All-time downloads
2.6K
Public
Parameters
3.5B
13.9 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors6.9 GB · 100%
From the Hugging Face model README
This repository contains a DFlash draft model for thinkingmachines/Inkling-NVFP4. It is not a standalone language model. It is intended to be paired with the target model in a speculative decoding server.
This is an early preview release, and the drafter is still training. It uses all causal sliding-window attention (SWA) layers.
DFlash uses a lightweight block diffusion draft model to propose multiple tokens in parallel. The target model verifies those proposals, improving serving throughput while preserving the target model's output distribution.
Here is an example deployment. Some options, including trtllm_mha for draft attention, are pending upstream SGLang support.
sglang serve \
--model-path thinkingmachines/Inkling-NVFP4 \
--tp 8 \
--trust-remote-code \
--quantization modelopt_fp4 \
--fp4-gemm-backend flashinfer_trtllm \
--moe-runner-backend flashinfer_trtllm_routed \
--enable-torch-symm-mem \
--attention-backend fa4 \
--mamba-radix-cache-strategy extra_buffer \
--disable-custom-all-reduce \
--page-size 128 \
--reasoning-parser inkling \
--tool-call-parser inkling \
--enable-multimodal \
--cuda-graph-backend-prefill breakable \
--enable-scattered-sconv \
--mem-fraction-static 0.78 \
--max-running-requests 32 \
--swa-full-tokens-ratio 0.10 \
--mamba-full-memory-ratio 0.10 \
--chunked-prefill-size 16384 \
--watchdog-timeout 900 \
--weight-loader-prefetch-checkpoints \
--weight-loader-prefetch-num-threads 8 \
--cuda-graph-max-bs-decode 32 \
--cuda-graph-max-bs-prefill 4096 \
--cuda-graph-bs-prefill 128 256 384 512 768 1024 1536 2048 2560 3072 3584 4096 \
--kv-cache-dtype mxfp8 \
--speculative-algorithm DFLASH \
--speculative-draft-model-path modal-labs/Inkling-DFlash \
--speculative-dflash-block-size 16 \
--speculative-draft-model-quantization fp8 \
--speculative-draft-attention-backend trtllm_mha \
--speculative-draft-kv-cache-dtype fp8_e4m3 \
--speculative-draft-window-size 4096 \
--host 0.0.0.0 \
--port 30000
This preview includes a preliminary accept-length evaluation at concurrency 1 and block size 16. A full benchmark suite, including throughput and higher-concurrency measurements, will follow.
thinkingmachines/Inkling-NVFP4mxfp8 KV cachefa4 target attention and trtllm_mha DFlash draft attentionfp8_e4m3 KV cache and a 4096-token draft windowcompletion_tokens / spec_verify_ct per generation turn, averaged across generation turnsMean DFlash accept length at concurrency 1.
| Workload | DFlash block=16 |
|---|---|
| GSM8K | 4.562 |
| MATH500 | 4.712 |
| HumanEval | 4.959 |
| MBPP | 3.907 |
| MT-Bench | 2.914 |
If you find DFlash useful, please cite the original paper:
@article{chen2026dflash,
title = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
journal = {arXiv preprint arXiv:2602.06036},
year = {2026}
}