Downloads · 30 days
373
49% of all-time downloads
Inferact/MiniMax-M3-EAGLE3-GQA-NVFP4
MiniMax-M3-EAGLE3-GQA-NVFP4 is a text generation model from Inferact. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
W4A4 NVFP4 MLP-quantized version of Inferact/MiniMax-M3-EAGLE3-GQA.
Downloads · 30 days
373
49% of all-time downloads
All-time downloads
767
Public
Parameters
2.9B
5.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.7 GB · 100%
How the weights are stored.
BF162.7B · 93%
From the Hugging Face model README
W4A4 NVFP4 MLP-quantized version of Inferact/MiniMax-M3-EAGLE3-GQA.
Mean accepted length and draft accept rate measured end-to-end against MiniMaxAI/MiniMax-M3-MXFP8 served with vLLM at tensor-parallel-size=4, num_speculative_tokens=3, greedy sampling (temperature=0, top_p=1.0), max-concurrency=16.
| Dataset | n | Mean accepted length | Draft accept rate | Per-position accept rate (pos 1 / 2 / 3) |
|---|---|---|---|---|
| MT-Bench | 64 | 2.663 | 55.42% | 0.742 / 0.534 / 0.386 |
| SPEED-Bench (qualitative) | 64 | 2.633 | 54.43% | 0.736 / 0.526 / 0.371 |