Downloads · 30 days
18
3% of all-time downloads
Sebesky/MiniMax-M3-EAGLE3-RTN-INT4
MiniMax-M3-EAGLE3-RTN-INT4 is a machine learning model from Sebesky. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
EAGLE3 speculative-decoding draft head for Sebesky/MiniMax-M3-W4A16-GPTQ: a single Llama-style decoder layer (hidden 6144, 64-head full MHA, headdim 128) with its own embedtokens and lmhead (vocab 200,064) — not share…
Downloads · 30 days
18
3% of all-time downloads
All-time downloads
526
Public
Parameters
3.3B
1.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.7 GB · 100%
How the weights are stored.
I323.3B · 99%
From the Hugging Face model README
EAGLE3 speculative-decoding draft head for
Sebesky/MiniMax-M3-W4A16-GPTQ:
a single Llama-style decoder layer (hidden 6144, 64-head full MHA, head_dim 128)
with its own embed_tokens and lm_head (vocab 200,064) — not shared with
the target. Everything is RTN-quantized to int4 (compressed-tensors
pack-quantized, group 128): Linear, embedding, and lm_head. ~1.6 GB.
--speculative-config '{"method":"eagle3","model":"Sebesky/MiniMax-M3-EAGLE3-RTN-INT4","num_speculative_tokens":3}'
Validated on 2x DGX Spark (GB10) at TP2 with draft_tensor_parallel_size: 2:
acceptance length 3.36 (78.6% draft acceptance) at 120k-token context,
decode 39.7 / 30.4 / 26.4 tok/s @512/65k/120k with 4-bit KV quantization.