Downloads · 30 days
24
3% of all-time downloads
OsaurusAI/MiniMax-M2.7-JANG_K
MiniMax-M2.7-JANG_K is a text generation model from OsaurusAI. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as other.
<p align="center"<img src="jangq-logo.png" width="160"/</p
Downloads · 30 days
24
3% of all-time downloads
All-time downloads
895
Public
Parameters
229B
85.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors85.9 GB · 100%
How the weights are stored.
U32229B · 100%
From the Hugging Face model README
MiniMax M2.7 — 86 GB on disk (down from ~230 GB FP8 source) — mixed-bit
JANG_K quantization using mx.quantize affine, prestacked switch_mlp.
mx.quantize, group_size=128):
down_proj: 4-bit (output enters residual stream — more sensitive)gate_proj: 2-bit + AWQ pre-scaling (gated activation)up_proj: 2-bit + AWQ pre-scaling (gated activation)q/k/v/o_proj: 8-bit affineblock_sparse_moe.switch_mlp.{gate,up,down}_proj of shape
(n_experts, out, in_packed) — instant cold load, no runtime sidecar.down_proj's output enters the residual stream and accumulates across
62 layers — quantization noise compounds. gate_proj and up_proj
enter through SwiGLU's multiplicative gate (silu(gate) × up) which
dampens noise. Spending 4 bits on down and 2 bits on gate/up gives
quality close to full-4-bit at considerably smaller size.
Activation-aware scaling on the 2-bit projections (gate_proj, up_proj):
(hidden,) scale: s = clip((max(|x|) + eps)^0.5, min=1.0)
(16 calibration prompts × ≤256 tokens; floor=1.0 prevents
inverse-fold from amplifying dead channels)W' = W * s[None, None, :]post_attention_layernorm.weight /= sdown_proj does not need AWQ — it stays at 4-bit.
Loadable via stock mlx-lm (no JANG runtime required):
from mlx_lm import load, generate
model, tok = load("JANGQ-AI/MiniMax-M2.7-JANG_K")
messages = [{"role": "user", "content": "What is the capital of France?"}]
prompt = tok.apply_chat_template(messages, add_generation_prompt=True,
tokenize=False)
print(generate(model, tok, prompt=prompt, max_tokens=128))
<think>\n after assistant prefix)qwen3 (extracts <think>...</think> blocks)minimaxprompt = tok.apply_chat_template(messages, add_generation_prompt=True,
tokenize=False, enable_thinking=False)
| Variant | Routed bits | Bundle size | Loader |
|---|---|---|---|
MiniMax-M2.7-JANGTQ | 2-bit codebook | 47 GB | jang_tools.load_jangtq |
MiniMax-M2.7-JANGTQ_K | mixed 2/4 codebook | 74 GB | jang_tools.load_jangtq |
MiniMax-M2.7-JANG_K (this) | mixed 2/4 affine + AWQ | 86 GB | stock mlx_lm |