Downloads · 30 days
140
8% of all-time downloads
OsaurusAI/MiniMax-M2.7-JANGTQ_K
MiniMax-M2.7-JANGTQ_K is a text generation model from OsaurusAI. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as other.
<p align="center"<img src="osaurus-x-banner.png" width="100%"/</p
Downloads · 30 days
140
8% of all-time downloads
All-time downloads
1.9K
Public
Parameters
22.8B
79.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors79.4 GB · 100%
How the weights are stored.
U3222.7B · 99%
From the Hugging Face model README
MiniMax M2.7 — 74 GB on disk (down from ~230 GB FP8 source) — mixed-bit JANGTQ_K quantization in JANGTQ-PRESTACK layout.
down_proj: 4-bit (output enters residual stream, more sensitive)gate_proj: 2-bit (gated activation, less sensitive)up_proj: 2-bit (gated activation)| Metric | Value | Setup |
|---|---|---|
| MMLU-200 | 93.5% (187/200) | thinking ON, q_per_subject=20, 10 subjects |
| Median speed | ~37 tok/s | M4 Max 128 GB, MLX 0.31 |
| GPU memory at load | ~75 GB | warm |
MMLU eval used the standard mmlu_jangtq_resume.py runner with the model's
default chat template (enable_thinking undefined → thinking ON, which the
M2.7 template auto-opens with <think>\n after the assistant prefix).
down_proj's output enters the residual stream and accumulates across
62 layers — quantization noise compounds. gate_proj and up_proj
enter through SwiGLU's multiplicative gate (silu(gate) × up) which
dampens noise. Spending 4 bits on down and 2 bits on gate/up gives
quality close to full-4-bit (~115 GB) at 64% the size.
| Variant | Routed bits (avg) | Size | MMLU-200 | Use case |
|---|---|---|---|---|
MiniMax-M2.7-JANGTQ | 2-bit | 56 GB | 91.5% | smallest, best for tight RAM |
MiniMax-M2.7-JANGTQ_K (this) | ~3-bit (mixed 2/4) | 74 GB | 93.5% | +2.0pp MMLU vs JANGTQ for +18 GB |
pip install jang-tools mlx-lm
from jang_tools.load_jangtq import load_jangtq_model
model, tokenizer = load_jangtq_model("OsaurusAI/MiniMax-M2.7-JANGTQ_K")
<think>\n after assistant prefix)messages = [{"role": "user", "content": "..."}]
inp = tokenizer.apply_chat_template(messages, add_generation_prompt=True, enable_thinking=False)
qwen3 (extracts <think>...</think> blocks)minimaxThe chat template ships with the enable_thinking switch correctly wired
both as a standalone chat_template.jinja AND inlined into
tokenizer_config.json["chat_template"] for engines that read inline
(vMLX, Swift swift-transformers).