Downloads · 30 days
27
13% of all-time downloads
dshive/MiMo-V2.5-Pro-FP4-DFlash
MiMo-V2.5-Pro-FP4-DFlash is a text generation model from dshive. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
27
13% of all-time downloads
All-time downloads
204
Public
Parameters
554B
570 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors570 GB · 100%
How the weights are stored.
U8531B · 96%
From the Hugging Face model README
<br/><br/>
<div align="center"> <picture> <source srcset="https://github.com/XiaomiMiMo/MiMo/raw/main/figures/Xiaomi_MiMo_darkmode.png?raw=true" media="(prefers-color-scheme: dark)"> <img src="https://github.com/XiaomiMiMo/MiMo/raw/main/figures/Xiaomi_MiMo.png?raw=true" width="60%" alt="Xiaomi-MiMo" /> </picture> </div> <br/> <div align="center" style="line-height: 1;"> <a href="https://huggingface.co/XiaomiMiMo" target="_blank">🤗 HuggingFace</a> | <a href="https://mimo.xiaomi.com/blog/mimo-tilert-1000tps" target="_blank">📰 Blog </a> <br/><a href="https://platform.xiaomimimo.com/ultraspeed" target="_blank">🎨 Xiaomi MiMo API Platform (Request Access) </a> | <a href="https://ultraspeed.xiaomimimo.com" target="_blank">🗨️ Xiaomi MiMo Studio (Free Trial) </a>
</div> <br/> <div align="center" style="line-height: 1.2;"> <strong>Community</strong><br/> <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro/blob/main/assets/wechat.jpg" target="_blank">WeChat Group</a> | <a href="https://discord.gg/kKC2kNnQEX" target="_blank">Discord</a> | <a href="https://t.me/+3T-I0pekOVIyNDBl" target="_blank">Telegram</a> | <a href="https://www.reddit.com/r/XiaomiMiMo_Official/" target="_blank">Reddit</a> </div> <br/>MiMo-V2.5-Pro-FP4-DFlash is the underlying model that powers MiMo-V2.5-Pro-UltraSpeed:
Together they cut both the per-parameter bit width and the number of backbone forward passes, the two dominant costs of trillion-parameter decoding.
At the trillion-parameter (1T) scale, even 8-bit (FP8/INT8) inference carries severe memory-footprint and memory-bandwidth costs. Lowering the parameter bit width translates directly into faster decoding. We therefore adopt FP4 quantization and block-diffusion speculative decoding. Key features of this release:
We quantize only the MoE experts to MXFP4 (block size 32) and keep attention projections and other modules at higher precision (the attention o_proj of every layer is excluded from FP4). With FP4 QAT, quality stays close to the FP8 baseline:
| Benchmark | MiMo-V2.5-Pro-FP8 | MiMo-V2.5-Pro-MXFP4 | Δ |
|---|---|---|---|
| General Agent | |||
| Claw-Eval (pass^3) | 63.8 | 67.8 | +6.27% |
| Humanity's Last Exam | 48.0 | 47.0 | -2.08% |
| Humanity's Last Exam (without tool) | 34.0 | 33.0 | -2.94% |
| Code Agent | |||
| SWE-Bench Pro | 57.2 | 58.8 | +2.80% |
| SWE-bench Verified | 78.9 | 77.4 | -1.90% |
Conventional speculative decoding relies on a small draft model to guess the next tokens, which the large model then verifies; the rejection-sampling verification keeps the output lossless. Its bottleneck is that draft quality bounds the acceptance rate, while a stronger draft costs more compute.
To break this trade-off we adopt the block-level masked parallel-prediction approach DFlash: the draft fills an entire block of masked positions in one forward pass. We landed this on MiMo-V2.5-Pro with custom optimizations for trillion-scale MoE and long-context serving, using the Muon second-order optimizer and model self-distillation so that even a small mask block keeps a strong acceptance rate while pushing the draft-stage cost close to its limit:
In practice, we further cap the mask block size at 8 to lower verification overhead and raise concurrency.
| Scenario | Acceptance Length |
|---|---|
| WebDev | 6.30 |
| Math500 | 5.56 |
| HumanEval | 4.54 |
| MT-Bench | 3.18 |
| SWE-Bench | 4.29 |
| Component | Backbone | DFlash Drafter |
|---|---|---|
| Architecture | MiMoV2ForCausalLM | DFlashDraftModel |
| Total / Active Params | 1.02T / 42B | 5-layer draft |
| Hidden Size | 6144 | 6144 |
| Num Layers | 70 | 5 |
| Num Attention Heads | 128 | 128 |
| Num KV Heads | 8 (GQA) | 8 (GQA) |
| Head Dim (QK / V) | 192 / 128 | 128 / 128 |
| SWA Window Size | 128 | 1024 |
| Block Size | — | 8 |
| Captured Backbone Layers | — | [0, 15, 31, 47, 69] |
| Backbone RoPE Base | 5,000,000 | 5,000,000 |
| Precision | MXFP4 (experts) Mixed | BF16 |
| Max Context Length | 1M | — |
DFlash inference with the FP4 backbone is supported in SGLang. The drafter is launched alongside the backbone via the speculative-decoding flags and inherits the backbone's tensor/expert-parallel topology.
The following is an example of running the model with SGLang. Point --model at this repository and --speculative-draft-model-path at its dflash/ subdirectory.
python3 -m sglang.launch_server \
--model MiMo-V2.5-Pro-FP4-DFlash \
--speculative-algorithm DFLASH \
--speculative-draft-model-path MiMo-V2.5-Pro-FP4-DFlash/dflash \
--speculative-num-draft-tokens 8 \
--ep-size 16 \
--tensor-parallel-size 16 \
--data-parallel-size 2 \
--enable-dp-attention \
--enable-dp-lm-head \
--quantization fp8 \
--attention-backend fa3 \
--moe-dense-tp-size 1 \
--dtype bfloat16 \
--mem-fraction-static 0.65 \
--context-length 65536 \
--page-size 1 \
--trust-remote-code \
--disable-overlap-schedule \
--skip-server-warmup \
--dist-init-addr ${MASTER_ADDR}:20000 \
--nnodes ${WORLD_SIZE} \
--node-rank ${RANK} \
--host 0.0.0.0 \
--port 29999
@misc{mimo2026v25pro_fp4dflash,
title={MiMo-V2.5-Pro-FP4-DFlash},
author={{Xiaomi MiMo Team}},
year={2026},
howpublished={\url{https://huggingface.co/collections/XiaomiMiMo/mimo-v25}},
}
For questions or feedback, reach us at [email protected] or join our community: