Downloads · 30 days
10
21% of all-time downloads
RESMP-DEV/MiMo-V2.5-ASR-FP8
MiMo-V2.5-ASR-FP8 is a automatic speech recognition model from RESMP-DEV. Use it when you need speech turned into text. The card lists the license as mit.
FP8-quantized build of XiaomiMiMo/MiMo-V2.5-ASR, the Xiaomi MiMo end-to-end ASR model with native Mandarin/English code-switching, Chinese dialects, song lyrics, noisy/multi-speaker robustness, and native punctuation.
Downloads · 30 days
10
21% of all-time downloads
All-time downloads
47
Public
Parameters
8B
8.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.7 GB · 100%
How the weights are stored.
F8_E4M37.4B · 92%
From the Hugging Face model README
FP8-quantized build of XiaomiMiMo/MiMo-V2.5-ASR, the Xiaomi MiMo end-to-end ASR model with native Mandarin/English code-switching, Chinese dialects, song lyrics, noisy/multi-speaker robustness, and native punctuation.
float8_e4m3fn, per-output-channel absmax scaling (one fp32 scale per row), baked at save time.torch._scaled_mm (FP8 tensor cores on Ada / Hopper / Blackwell).This roughly halves the LLM weight footprint (~32 GB bf16 → ~16-17 GB on disk).
from_pretrained checkpointmodel.safetensors stores custom FP8Linear buffers (*.weight_fp8, *.weight_scale),
not standard HF Linear weights. It must be loaded through the matching FP8Linear
modules. Use the loader below.
git clone https://github.com/XiaomiMiMo/MiMo-V2.5-ASR.git
cd MiMo-V2.5-ASR
pip install -r requirements.txt
pip install flash-attn==2.7.4.post1 # required by the audio tokenizer
hf download XiaomiMiMo/MiMo-Audio-Tokenizer --local-dir ./models/MiMo-Audio-Tokenizer
hf download Infatoshi/MiMo-V2.5-ASR-FP8 --local-dir ./MiMo-V2.5-ASR-FP8
Then install the maintained loader from
RESMP-DEV/mimo-asr-fp8 and load
with its FP8Linear implementation:
from quantize_fp8 import load_fp8_model
mimo = load_fp8_model(
fp8_dir="./MiMo-V2.5-ASR-FP8",
tokenizer_path="./models/MiMo-Audio-Tokenizer",
repo_root=".", # the cloned MiMo-V2.5-ASR repo
)
print(mimo.asr_sft("audio.wav", audio_tag="<english>"))
Per-output-channel absmax dequant error vs the original fp32 weights, sampled across depth (layers 0/17/35), all attn+mlp projections, lm_head, and the audio local transformer:
This is the expected magnitude for fp8 e4m3 with per-channel scaling (3 mantissa bits).
cu128). torch 2.6 cu124 ships no sm_120 kernels and will fail with
"no kernel image is available for execution on the device".Derivative of an MIT-licensed model; original credit to the Xiaomi MiMo team.