Downloads · 30 days
112
7% of all-time downloads
majentik/Mistral-Small-4-119B-RotorQuant-MLX-8bit
Mistral-Small-4-119B-RotorQuant-MLX-8bit is a text generation model from majentik. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
[!TIP] KV-cache quantization without any fork (recommended, 2026): upstream llama.cpp/Ollama now cover this natively — use -ctk q80 -ctv q80 (~half KV memory, negligible quality loss: perplexity +0.002–0.05) or -ctk q…
Downloads · 30 days
112
7% of all-time downloads
All-time downloads
1.7K
Public
Parameters
119B
127 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors127 GB · 100%
How the weights are stored.
U32119B · 100%
From the Hugging Face model README
<!-- kv-upstream-note -->[!TIP] KV-cache quantization without any fork (recommended, 2026): upstream llama.cpp/Ollama now cover this natively — use
-ctk q8_0 -ctv q8_0(~half KV memory, negligible quality loss: perplexity +0.002–0.05) or-ctk q4_0 -ctv q4_0(~quarter memory, ≈7.6% perplexity increase). In Ollama:OLLAMA_KV_CACHE_TYPE=q8_0withOLLAMA_FLASH_ATTENTION=1. Keep K and V types symmetric to stay on the fast fused Flash-Attention path. Since April 2026, mainline llama.cpp also applies Hadamard rotation to KV activations (PR #21038), which greatly improves low-bit KV quality (opt-out:LLAMA_ATTN_ROT_DISABLE=1).The RotorQuant/TurboQuant fork flow below is experimental/legacy: the TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork is unmaintained relative to mainline. It is NOT required to use this model.
Dual compression: 8-bit MLX weight quantization + RotorQuant KV cache quantization for Mistral Small 4 on Apple Silicon.
This repository provides an 8-bit weight-quantized MLX conversion of mistralai/Mistral-Small-4-119B-2603 with RotorQuant KV cache quantization support. Designed for efficient inference on Apple Silicon Macs with excellent throughput.
Approximate model size: ~120 GB
This model applies two complementary compression techniques:
Together, these make it feasible to run a 119B-parameter MoE model on high-memory Apple Silicon machines with excellent throughput.
| Property | Value |
|---|---|
| Base Model | Mistral Small 4 (March 2026) |
| Total Parameters | 119B |
| Active Parameters | 6.5B per token (Sparse MoE) |
| Architecture | Sparse MoE -- 128 experts, 4 active per token |
| Context Length | 256K tokens |
| Modality | Text + Images (multimodal) |
| Capabilities | Thinking / reasoning, tool use, multilingual |
| License | Apache 2.0 |
| Weight Quantization | 8-bit (MLX) |
| KV Cache Quantization | RotorQuant 3-bit |
| Configuration | Weights | KV Cache (256K) | Total |
|---|---|---|---|
| FP16 baseline | ~238 GB | ~32 GB | ~270 GB |
| This model (8-bit MLX + RotorQuant) | ~120 GB | ~6.5 GB | ~126.5 GB |
Note: This is a Sparse MoE model -- only 6.5B parameters are active per token, so inference is fast despite the 119B total parameter count.
from mlx_lm import load, generate
model, tokenizer = load("majentik/Mistral-Small-4-119B-RotorQuant-MLX-8bit")
prompt = "Explain sparse mixture-of-experts architectures."
messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
response = generate(model, tokenizer, prompt=text, max_tokens=512)
print(response)
RotorQuant and TurboQuant are this project's release labels, not distinct
quantization algorithms — for any given tier, both brand repos carry
byte-identical weights produced with the standard MLX / llama.cpp quantizers.
No brand-specific speedup is claimed or measured. The KV-cache fork these
labels originally referred to is legacy; for KV-cache memory savings use the
upstream options described above (-ctk/-ctv q8_0, OLLAMA_KV_CACHE_TYPE).
| Method | Prefill Speed | Decode Speed | Memory Savings | Reference |
|---|---|---|---|---|
| TurboQuant | Baseline | Baseline | High | arXiv: 2504.19874 |
This model requires approximately 127 GB total memory at 256K context. Recommended hardware: