Downloads · 30 days
134
7% of all-time downloads
avlp12/MiniMax-M3-Alis-MLX-Dynamic
MiniMax-M3-Alis-MLX-Dynamic is a image-text-to-text model from avlp12. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as other.
The first MLX quantization of MiniMax-M3 that keeps the full vision-language model. Sensitivity-graded mixed-precision builds of MiniMaxAI/MiniMax-M3 (427B total / ~23B active MoE VL, MiniMax Sparse Attention, 1M cont…
Downloads · 30 days
134
7% of all-time downloads
All-time downloads
1.9K
Public
Parameters
427B
1.2 TB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors193 GB · 100%
How the weights are stored.
U32426B · 100%
From the Hugging Face model README
The first MLX quantization of MiniMax-M3 that keeps the full vision-language model. Sensitivity-graded mixed-precision builds of MiniMaxAI/MiniMax-M3 (427B total / ~23B active MoE VL, MiniMax Sparse Attention, 1M context) for Apple Silicon, sized for 512 GiB and 256 GiB machines.
Exact parameter count, measured across all 59 bf16 source shards: 427.04B (language model 426.18B + vision tower 0.63B + patch-merge 0.19B + projector 0.05B).
Existing MLX quants of M3 are text-only extractions — the vision tower, multimodal
projector and patch-merge MLP are deleted. These builds keep all of it: image and
video prompts work end-to-end through mlx-vlm.
Default branches load on stock mlx-vlm and on oMLX — no fork, no patch.
| branch | routed experts | shared expert | attn/dense | embed / head | vision | size | fits |
|---|---|---|---|---|---|---|---|
main (T256) | 3-bit g64 | 3-bit (packed) | 6-bit | 6b / 8b | bf16 | 192.6 GB (3.65 bpw) | 256 GiB Macs, no sysctl needed |
t512 (T512) | 6-bit g64 | 6-bit (packed) | 8-bit | 8b / 8b | bf16 | 350.8 GB (6.57 bpw) | 512 GiB Macs |
t512ref (T512REF) | 8-bit g64 | 8-bit | 8-bit | 8b / 8b | bf16 | 454.9 GB (8.52 bpw) | 512 GiB Macs (max quality / reference) |
All three load on stock mlx-vlm and oMLX with no patch (t512ref keeps the unpacked layout
but is uniform 8-bit, so its shared expert is already 8-bit and concatenates cleanly).
Never quantized, in every build:
e_score_correction_bias — fp32 (discrete top-4 expert selection)index_q_proj/index_k_proj) — bf16 (they pick the top-16
attention blocks; a flipped selection reads different history, so this control path stays exact)M3's always-on shared expert sees 100% of tokens (a routed expert sees ~3%), so holding
it at 8-bit while the routed bank drops to 3/6-bit is a real quality-per-GB lever (+0.4% size).
But MLX stores it packed into the same 129-wide SwitchLinear as the routed experts, and a
packed bank can hold only one bit-width — so an 8-bit shared expert over a 3/6-bit routed
bank needs the unpacked layout, which stock mlx-vlm and oMLX cannot load (they unconditionally
re-pack and can't concatenate mixed bit-widths — see
discussion #1).
These builds therefore keep the shared expert packed at routed bits so they load
everywhere with no fork. If you want the 8-bit-shared variant, rebuild it with
pack_shared_expert=false on the patched
mlx-vlm — it is not hosted here.
Stock, no fork:
pip install mlx-vlm # >= 0.6.5
# text
mlx_vlm.generate --model avlp12/MiniMax-M3-Alis-MLX-Dynamic \
--prompt "Explain MoE routing in three sentences." --max-tokens 300
# image
mlx_vlm.generate --model avlp12/MiniMax-M3-Alis-MLX-Dynamic \
--image photo.jpg --prompt "Describe this image." --max-tokens 300
# video
mlx_vlm.generate --model avlp12/MiniMax-M3-Alis-MLX-Dynamic \
--video clip.mp4 --prompt "What happens in this clip?" --max-tokens 300
Python:
import mlx.core as mx
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
model, processor = load("avlp12/MiniMax-M3-Alis-MLX-Dynamic") # or revision="t512"
mx.set_wired_limit(mx.metal.device_info()["max_recommended_working_set_size"])
prompt = apply_chat_template(processor, model.config, "Describe this image.", num_images=1)
out = generate(model, processor, prompt, image=["photo.jpg"], max_tokens=300)
print(out.text)
Sampling: MiniMax recommends temperature=1.0, top_p=0.95. The model thinks in
<mm:think>…</mm:think> before answering; budget max_tokens accordingly.
Same-machine (M3 Ultra 512 GB), deterministic contexts (EN/KO/code/econ prose), KL on the reference's top-256 support; long-context on a 16K document exercising the sparse-attention path (MSA only activates beyond ~2.2K tokens — short evals cannot see it).
These numbers were measured on an earlier 8-bit-shared-expert prototype. The published packed builds keep the shared expert at routed bits (T256 → 3-bit, T512 → 6-bit): the T512 delta is negligible (routed already 6-bit), the T256 delta is bounded (one always-on FFN among a 3-bit bank). Figures are indicative; the qualitative ranking holds.
| metric | T512 | T256 (main) |
|---|---|---|
| KL vs REF, short (512 tok) | 0.0184 nats | 0.1243 nats |
| top-1 agreement, short | 97.4% | 90.4% |
| KL vs REF @16K | 0.0087 nats | 0.0011 nats |
| top-1 agreement @16K | 99.2% | 100.0% |
| NIAH @16K (3 depths) | 3/3 | 3/3 |
| sparse block-selection overlap vs REF @16K | 54% | 48% |
| vision (figure description + OCR of axis labels/annotations) | pass | pass |
| decode tok/s (short / 2.4K ctx) | 22.5 / 17.7 | 28.1 / 20.6 |
| prefill tok/s @2.4K | 344 | 355 |
| peak memory @2.4K | 362 GB | 205 GB |
Reference itself: PPL 2.851 on the eval slice (T512 2.900, T256 2.994), NIAH 3/3, decode 20.6 tok/s.
Notes, honestly stated:
KV cache ≈ 134 KiB/token fp16 (60L × 4 KV heads × 128d K+V, + 57 index-key caches) → ~13.7 GB per 100K tokens.
| machine | default GPU limit (~75% RAM) | with sudo sysctl iogpu.wired_limit_mb=... |
|---|---|---|
| 512 GiB | ~412 GB — T512 fits, REF needs bump | up to ~475 GB: REF + ~20 GB ctx, T512 + 1M-token ctx |
| 256 GiB | ~206 GB — T256 fits stock, ~11 GB ctx (~80K tok) | up to ~240 GB: ~300K+ tok |
A 128 GiB build was evaluated and cut: the routed-expert floor alone (2-bit g128 ≈ 117 GB) plus essentials exceeds a 128 GiB Mac's realistic wired ceiling — it cannot load, so we won't ship it.
mlx_vlm.convert + a custom per-module quant predicate.