Downloads · 30 days
22
2% of all-time downloads
majentik/gemma-4-12B-RotorQuant-MLX-8bit
gemma-4-12B-RotorQuant-MLX-8bit is a image-text-to-text model from majentik. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as gemma.
google/gemma-4-12B @ 023679ed352de9bb66cc873c9009ce3482585c08 quantized pack, published as majentik/gemma-4-12B-RotorQuant-MLX-8bit.
Downloads · 30 days
22
2% of all-time downloads
All-time downloads
1.2K
Public
Parameters
12B
12.7 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors12.7 GB · 100%
How the weights are stored.
U3212B · 100%
From the Hugging Face model README
google/gemma-4-12B @ 023679ed352de9bb66cc873c9009ce3482585c08 quantized pack, published as majentik/gemma-4-12B-RotorQuant-MLX-8bit.
MLX quantization via mlx-vlm 0.6.3 (8bit, group_size 64); vision + audio towers retained in BF16 (not quantized).
Released under the RotorQuant line. RotorQuant and TurboQuant are this project's release labels for this pack, not distinct quantization algorithms — both brand repos for a given tier carry byte-identical weights, produced once and published under two names. No brand-specific speedup is claimed or measured for either label.
This pack is image-text-to-text capable: the vision and audio towers ship in BF16 alongside the quantized text tower, so image (and audio) inputs are supported end to end via mlx-vlm.
Base-model note: this is the raw completion model (not instruction-tuned). It is not instruction-aligned for multimodal Q&A — expect free-form continuation behavior rather than chat-style image description, even though the vision/audio towers are present.
mlx-vlm suppress-tokens note: the base model's generation_config.json carries suppress_tokens for six multimodal placeholder token ids so that transformers generation is clean. mlx-vlm does not honor suppress_tokens automatically, so text-only generation with mlx-vlm can leak placeholder tokens (observed as a degenerate A<image|>A<image|>... loop in smoke testing) unless you suppress them yourself:
from mlx_vlm import load, generate
model, processor = load("majentik/gemma-4-12B-RotorQuant-MLX-8bit")
suppress = {255999: -1e9, 256000: -1e9, 258880: -1e9, 258881: -1e9, 258882: -1e9, 258883: -1e9}
output = generate(
model, processor, prompt,
max_tokens=256,
logit_bias=suppress,
)
Governed by the Gemma Terms of Use. See the upstream repo for the full license text.