Downloads · 30 days
0
Reza2kn/mega-asr-nvfp4
mega-asr-nvfp4 is a automatic speech recognition model from Reza2kn. Use it when you need speech turned into text. It is set up for modelopt. The card lists the license as apache-2.0.
NVFP4 (4-bit floating-point: E2M1 mantissa with per-block FP8 scaling) deployment of the LLM portion of zhifeixie/Mega-ASR, quantized via NVIDIA Model Optimizer
Downloads · 30 days
0
Access
Public
Updated Sep 9, 2026
Repo size
4.7 GB
Likes
1
Public
Click a slice to open those files.
.safetensors3.4 GB · 73%
From the Hugging Face model README
NVFP4
(4-bit floating-point: E2M1 mantissa with per-block FP8 scaling) deployment of
the LLM portion of zhifeixie/Mega-ASR,
quantized via NVIDIA Model Optimizer
(nvidia-modelopt) with the NVFP4_AWQ_LITE_CFG activation-aware recipe.
Targets RTX 50-series (Blackwell) for native NVFP4 GEMM acceleration. Earlier Ada/Hopper GPUs can run the same checkpoint via the modelopt fake-quant simulation (correctness preserved, no perf win without Blackwell NVFP4 tensor cores).
| File | Size | Role |
|---|---|---|
nvfp4/model.safetensors | 3.44 GB | Qwen3 1.7B LLM, NVFP4 weights + AWQ-Lite scaling factors. Saved via modelopt.torch.save_pretrained — the weights are stored in their original bf16 layout alongside the quantization scale tensors; the runtime packs them to NVFP4 on first forward. |
nvfp4/config.json + tokenizer/* | — | HF config + Qwen3-ASR tokenizer (with <|audio_pad|>, <asr_text>, etc.) |
onnx/audio_encoder_fp32.onnx | 1.27 GB | 24-layer Whisper-style audio encoder (ONNX fp32, run via onnxruntime; NVFP4 port not done — the encoder is small enough that it doesn't benefit much from FP4) |
examples/*.wav | ~3 MB | 8 noisy benchmark clips from Voices-in-the-Wild-Bench |
nvfp4_quantize.py | — | The PTQ script (modelopt forward-loop calibration) |
inference_bench.py | — | End-to-end ASR pipeline + 8-clip VITW bench |
8-clip Voices-in-the-Wild-Bench
agreement (1 − WER), prompt forced to language English, run on the RTX 5080
Laptop (Blackwell, compute_cap 12.0). Same ONNX fp32 audio encoder as the
other backends:
| Per-sample | NVFP4 (this repo) | ONNX GPTQ | MLX mixed 8/4 | CoreML mixed 8/4 |
|---|---|---|---|---|
| distortion | 100% | 100% | 100% | 100% |
| dropout | 100% | 100% | 100% | 100% |
| echo (hard, reverb) | 64.7% | 82.4% | 64.7% | 64.7% |
| far_field | 100% | 100% | 100% | 100% |
| mixed | 100% | 100% | 100% | 100% |
| noise | 100% | 100% | 100% | 100% |
| obstructed | 100% | 100% | 94.1% | 100% |
| recording (hard, truncated) | 66.7% | 60.0% | 60.0% | 60.0% |
| AVERAGE | 91.4% | 92.7% | 92.2% | 90.6% |
Notable: NVFP4 ties or beats the others on every clean sample, wins on
recording by 6.7 pts (66.7% vs 60% everywhere else — the AWQ-Lite
activation-aware scaling helped recover the truncated-audio decode), and
ties MLX/CoreML on echo. The 1.3% gap to ONNX GPTQ is entirely on echo
(64.7% vs 82.4%) where GPTQ's per-column Hessian-based error redistribution
captures something AWQ-Lite's per-channel scaling doesn't.
The AWQ-Lite variant runs an extra pass that computes a per-channel activation magnitude and rescales weights vs. activations to put more of the dynamic range into the "important" channels (channels with large activation amplitudes) before applying NVFP4 — net effect is recovering some quality lost to the E2M1 grid.
pip install nvidia-modelopt transformers safetensors torch onnxruntime soundfile librosa
git clone https://huggingface.co/Reza2kn/mega-asr-nvfp4
cd mega-asr-nvfp4
python inference_bench.py \
--model nvfp4 \
--encoder onnx/audio_encoder_fp32.onnx \
--examples-dir examples \
--qwen-asr-dir <Qwen3-ASR-1.7B HF dir> \
--skip-quant # weights already quantized
# Convert HF checkpoint → TensorRT-LLM checkpoint
python -m tensorrt_llm.examples.qwen.convert_checkpoint \
--model_dir nvfp4 --output_dir trtllm_ckpt \
--dtype bfloat16 --use_fp4
# Build engine
trtllm-build --checkpoint_dir trtllm_ckpt --output_dir trtllm_engine \
--gemm_plugin fp4 --max_input_len 512 --max_seq_len 600
(The TRT-LLM engine path is on the roadmap; this repo currently ships the modelopt-saved HF checkpoint, which runs as fake-quant on any GPU.)
import modelopt.torch.quantization as mtq
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("Qwen3-ASR-1.7B-LLM",
torch_dtype=torch.bfloat16,
device_map="cuda")
# Calibration with 168 English VITW samples (audio embeds scattered at
# <|audio_pad|> positions — same set used for the ONNX GPTQ release)
calib_batches = build_calibration_batches(...)
def forward_loop(m):
for b in calib_batches:
with torch.no_grad():
m(**b)
mtq.quantize(model, mtq.NVFP4_AWQ_LITE_CFG, forward_loop)
model.save_pretrained("nvfp4")
168 calibration batches, ~3 min on the RTX 5080. The AWQ-Lite recipe does two forward passes per batch — one for activation magnitude estimation, one for the actual quantization apply step — explaining the doubled count in the log.
This model distribution is licensed under the Apache License, Version 2.0. See LICENSE. Existing third-party copyright, license, and attribution notices remain applicable.
Upstream: Qwen/Qwen3-ASR-1.7B; declared license: apache-2.0.
Upstream: zhifeixie/Mega-ASR; declared license: apache-2.0.