Downloads · 30 days
28
3% of all-time downloads
Ex0bit/MiniMax-SLURPY-DQ-MLX
MiniMax-SLURPY-DQ-MLX is a text generation model from Ex0bit. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as other.
Downloads · 30 days
28
3% of all-time downloads
All-time downloads
816
Public
Parameters
229B
72.7 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors72.7 GB · 100%
How the weights are stored.
U32229B · 100%
From the Hugging Face model README

Per-tensor mixed-precision quantization of MiniMax-SLURPY for Apple Silicon — 2.54 BPW with 498 per-tensor-projection allocations (plus 16,122 per-expert PRISM decisions collapsed into MLX's SwitchGLU format).
The full SLURPY model (228.7B params) compressed from 215 GB → 68 GB (68% reduction) using PRISM Dynamic Quantization — a per-tensor-class mixed-precision allocation derived entirely from weight structure sensitivity analysis. Zero calibration data, zero training, zero datasets.
Created by Ex0bit
| Property | Value |
|---|---|
| Base Model | Ex0bit/MiniMax-SLURPY |
| Architecture | MiniMax M2 MoE (256 experts, top-8) |
| Parameters | 228.7B total / ~10B active |
| Quantization | PRISM-DYNAMIC-QUANT (MLX native) |
| Achieved BPW | 2.54 |
| File Size | 68 GB (vs 215 GB source = 68% reduction) |
| Per-tensor overrides | 498 (MoE: per-layer-projection modal of 16,122 per-expert decisions) |
| Default precision | 2-bit |
| Group size | 64 |
| Context Length | 196,608 tokens |
| Runtime | mlx-lm (Apple Silicon Metal) |
| Creator | Ex0bit |
A mathematically unique Designer Baby of MiniMax-M2.5 and MiniMax-M2.7 — neither parent, entirely its own model.
SLURPY inherits M2.5's architect-first coding style and MIT freedom, absorbs M2.7's RL-tuned precision on multi-agent collaboration and real-world engineering — without a single training step.
| Benchmark | M2.5 | M2.7 | SLURPY |
|---|---|---|---|
| HumanEval pass@5 | 85.4% | — | 89.6% |
| SWE-Bench Verified | 80.2% | — | inherited |
| SWE-Pro | — | 56.2% | inherited |
| MLE Bench Lite | — | 66.6% | inherited |
| GDPval-AA ELO | — | 1495 | inherited |
See Ex0bit/MiniMax-SLURPY for full benchmark details.
This model uses PRISM Dynamic Quantization — a per-tensor mixed-precision allocation that assigns different quantization types to different tensor classes based on weight structure sensitivity analysis.
Unlike uniform quantization (Q3, Q4, Q5), PRISM-DQ analyzes each tensor's sensitivity to quantization error and allocates precision where it matters most. Critical tensors (attention projections, key MoE experts, lm_head) receive higher precision while less impactful tensors get aggressive compression.
PRISM produced 16,122 per-expert decisions (256 experts × 62 layers × 3 projections, plus attention and embeddings). MLX's SwitchGLU packs all 256 experts per layer-projection into a single 3D tensor sharing one bit width, so the per-expert decisions collapse to the modal bit width for each of the 186 MoE projections. The remaining 312 per-tensor decisions (attention, embeddings, lm_head, routers) retain full PRISM granularity, giving 498 effective overrides.
The model's config.json contains the per-tensor quantization overrides that mlx-lm loads natively — no custom runtime required. Apple Silicon's compiled Metal kernels automatically handle mixed-precision tensors in a single forward pass at full GPU speed.
No calibration data, no importance matrices, no training data required.
Identical to MiniMax-M2.5 / M2.7 — quantization-only:
minimax_m2 / MiniMaxM2ForCausalLM<think>...</think> (always-on)trust_remote_code=True requiredpip install mlx-lm
# Interactive chat
mlx_lm.chat --model Ex0bit/MiniMax-SLURPY-PRISM-3BPW-MLX \
--temperature 1.0 --top-p 0.95 --max-tokens 4096
# Single prompt
python -m mlx_lm.generate \
--model Ex0bit/MiniMax-SLURPY-PRISM-3BPW-MLX \
--prompt "Write a Python function that reverses a linked list." \
--max-tokens 2048 \
--temp 1.0 --top-p 0.95
from mlx_lm import load, generate
model, tokenizer = load("Ex0bit/MiniMax-SLURPY-PRISM-3BPW-MLX")
response = generate(
model, tokenizer,
prompt="Write a Python function that reverses a linked list.",
max_tokens=2048,
temp=1.0,
top_p=0.95,
)
print(response)
| Parameter | Value |
|---|---|
| temperature | 1.0 |
| top_p | 0.95 |
| top_k | 40 |
MiniMax-M2 uses interleaved thinking. The model outputs <think>...</think> blocks during generation. You must pass these back verbatim in conversation history. Removing them degrades performance.
Same format as base SLURPY. Tool calls use <minimax:tool_call> / </minimax:tool_call> XML wrappers:
<minimax:tool_call>
<invoke name="get_weather">
<parameter name="city">San Francisco</parameter>
</invoke>
</minimax:tool_call>
For non-Apple platforms, use the FP8 Ex0bit/MiniMax-SLURPY variant with vLLM.
config.json with 498 per-tensor quantization overrides (collapsed from 16,122 PRISM decisions via SwitchGLU packing)chat_template.jinja — M2.7's chat template with tool calling supportmodeling_minimax_m2.py / configuration_minimax_m2.py — custom model code (inherited from base)Modified MIT — same as MiniMax-M2.5. See LICENSE for full text.
The only modification to the standard MIT license: if the Software (or any derivative works) is used for commercial products or services with more than 100 million monthly active users or more than $30M annual recurring revenue, you must prominently display "MiniMax M2" on the user interface.
@misc{minimax-slurpy-prism-mlx-2026,
title={MiniMax-SLURPY-PRISM-3BPW-MLX: Per-tensor mixed-precision quantization of MiniMax-SLURPY for Apple Silicon},
author={Ex0bit},
year={2026},
url={https://huggingface.co/Ex0bit/MiniMax-SLURPY-PRISM-3BPW-MLX}
}