Downloads · 30 days
90
8% of all-time downloads
majentik/gemma-4-E4B-RotorQuant-MLX-2bit
gemma-4-E4B-RotorQuant-MLX-2bit is a image-text-to-text model from majentik. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as apache-2.0.
[!TIP] KV-cache quantization without any fork (recommended, 2026): upstream llama.cpp/Ollama now cover this natively — use -ctk q80 -ctv q80 (~half KV memory, negligible quality loss: perplexity +0.002–0.05) or -ctk q…
Downloads · 30 days
90
8% of all-time downloads
All-time downloads
1.1K
Public
Parameters
8B
3.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.3 GB · 99%
How the weights are stored.
U327.5B · 94%
From the Hugging Face model README
<!-- kv-upstream-note -->[!TIP] KV-cache quantization without any fork (recommended, 2026): upstream llama.cpp/Ollama now cover this natively — use
-ctk q8_0 -ctv q8_0(~half KV memory, negligible quality loss: perplexity +0.002–0.05) or-ctk q4_0 -ctv q4_0(~quarter memory, ≈7.6% perplexity increase). In Ollama:OLLAMA_KV_CACHE_TYPE=q8_0withOLLAMA_FLASH_ATTENTION=1. Keep K and V types symmetric to stay on the fast fused Flash-Attention path. Since April 2026, mainline llama.cpp also applies Hadamard rotation to KV activations (PR #21038), which greatly improves low-bit KV quality (opt-out:LLAMA_ATTN_ROT_DISABLE=1).The RotorQuant/TurboQuant fork flow below is experimental/legacy: the TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork is unmaintained relative to mainline. It is NOT required to use this model.
2-bit weight-quantized MLX version of google/gemma-4-E4B with the legacy RotorQuant KV-cache fork (superseded by upstream llama.cpp KV options). Optimized for Apple Silicon inference via the MLX framework. The most aggressive quantization, fitting the full model in the smallest possible footprint.
Approximate model size: ~1.2 GB
| Property | Value |
|---|---|
| Base Model | google/gemma-4-E4B |
| Parameters | ~4 billion |
| Architecture | Dense transformer |
| Modality | Multimodal: image + text input, text output |
| License | Apache 2.0 |
| Weight Quantization | 2-bit (~1.2 GB) |
| KV-Cache Quantization | RotorQuant |
| Framework | MLX (Apple Silicon) |
import mlx.core as mx
from mlx_lm import load, generate
model, tokenizer = load("majentik/gemma-4-E4B-RotorQuant-MLX-2bit")
prompt = "The history of artificial intelligence began"
response = generate(model, tokenizer, prompt=prompt, max_tokens=512)
print(response)
For multimodal usage with images:
from mlx_vlm import load, generate
model, processor = load("majentik/gemma-4-E4B-RotorQuant-MLX-2bit")
prompt = "Describe the contents of this image."
output = generate(model, processor, prompt=prompt, image="path/to/image.jpg", max_tokens=512)
print(output)
RotorQuant and TurboQuant are this project's release labels, not distinct
quantization algorithms — for any given tier, both brand repos carry
byte-identical weights produced with the standard MLX / llama.cpp quantizers.
No brand-specific speedup is claimed or measured. The KV-cache fork these
labels originally referred to is legacy; for KV-cache memory savings use the
upstream options described above (-ctk/-ctv q8_0, OLLAMA_KV_CACHE_TYPE).
| Method | Prefill Speed | Decode Speed | Memory Savings | Reference |
|---|---|---|---|---|
| TurboQuant | 1x (baseline) | 1x (baseline) | High | arXiv: 2504.19874 |
| Precision | Approximate Size | MLX Variant |
|---|---|---|
| FP16 (original) | ~8 GB | -- |
| 8-bit quantized | ~4 GB | RotorQuant-MLX-8bit |
| 4-bit quantized | ~2.3 GB | RotorQuant-MLX-4bit |
| 2-bit quantized | ~1.2 GB | This model |
This model requires approximately 1.2 GB of unified memory. Recommended hardware:
| Bits | Approx size | Use case | Recommendation |
|---|---|---|---|
| 2-bit | ~1.0 GB | Aggressive quantization | Very low-RAM Macs |
| 3-bit | ~1.4 GB | Lossy but small | Low-RAM Macs |
| 4-bit | ~1.7 GB | Balanced default | Recommended for most Macs |
| 5-bit | ~2.0 GB | Higher fidelity | Quality-sensitive |
| 6-bit | ~2.4 GB | Approaching FP16 quality | High-fidelity |
| 8-bit | ~3.0 GB | Near-lossless reference | Fidelity-critical work |
(Current variant — 2bit — is bolded.)
(Showing 14 sibling variants under majentik/gemma-4-e4b-*. The current variant — RotorQuant-MLX-2bit — is bolded.)
| Variant | Runtime | Approx size | Use case |
|---|---|---|---|
| RotorQuant-GGUF-IQ4_XS | llama.cpp | ~3.4 GB | Lossy 4-bit, low-RAM CPU/edge |
| RotorQuant-GGUF-Q2_K | llama.cpp | ~2.4 GB | Lossy, low-RAM CPU/edge |
| RotorQuant-GGUF-Q3_K_M | llama.cpp | ~3.1 GB | Smaller 3-bit, CPU-friendly |
| RotorQuant-GGUF-Q4_K_M | llama.cpp | ~4.4 GB | Balanced default |
| RotorQuant-GGUF-Q5_K_M | llama.cpp | ~5.3 GB | Higher fidelity, more RAM |
| RotorQuant-GGUF-Q8_0 | llama.cpp | ~8.4 GB | Near-lossless reference |
| RotorQuant-MLX-2bit | mlx-lm | ~1.3 GB | Apple Silicon, smallest |
| RotorQuant-MLX-4bit | mlx-lm | ~2.5 GB | Apple Silicon balanced |
| RotorQuant-MLX-8bit | mlx-lm | ~4.7 GB | Apple Silicon reference |
| TurboQuant | runtime modifier | n/a | KV-cache root (weight-agnostic) |