Downloads · 30 days
47
35% of all-time downloads
ToPo-ToPo/diffusiongemma-26B-A4B-it-mlx-4bit
diffusiongemma-26B-A4B-it-mlx-4bit is a image-text-to-text model from ToPo-ToPo. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as apache-2.0.
MLX 4bit conversion of google/diffusiongemma-26B-A4B-it (mlx-vlm). Block-diffusion LM built on Gemma 4 (25.2B total / 3.8B active, MoE 128+1 experts, vision).
Downloads · 30 days
47
35% of all-time downloads
All-time downloads
134
Public
Parameters
25.8B
16.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.5 GB · 100%
How the weights are stored.
U3225.3B · 98%
From the Hugging Face model README
MLX 4bit conversion of google/diffusiongemma-26B-A4B-it (mlx-vlm).
Block-diffusion LM built on Gemma 4 (25.2B total / 3.8B active, MoE 128+1 experts, vision).
google/diffusiongemma-26B-A4B-it (license: apache-2.0)mlx_vlm.convert (4bit affine, group_size=64), ~5.130 bpwmodel_type: diffusion_gemma is supported natively by mlx-vlm 0.6.9; no patch needed.chat_template.jinja is not the base repo's copy: it is the patched Gemma 4 Canonical
Chat Template from ToPo-ToPo/gemma-4-26B-A4B-it-mlx-4bit,
which suppresses the thinking channel when enable_thinking is false (otherwise the literal
word thought leaks into the answer). Thinking is off by default. Weights are unaffected —
restore the base repo's template for stock behaviour.from mlx_vlm import load
model, processor = load("ToPo-ToPo/diffusiongemma-26B-A4B-it-mlx-4bit")
Diffusion generation takes its own flags:
python -m mlx_vlm generate --model ToPo-ToPo/diffusiongemma-26B-A4B-it-mlx-4bit \
--prompt "Why is the sky blue?" \
--max-tokens 256 --max-denoising-steps 48 --diffusion-sampler entropy-bound