Downloads · 30 days
11
4% of all-time downloads
ToPo-ToPo/gemma-4-31b-it-mlx-4bit
gemma-4-31b-it-mlx-4bit is a image-text-to-text model from ToPo-ToPo. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as gemma.
MLX 4bit conversion of google/gemma-4-31b-it for Apple Silicon (mlx-vlm).
Downloads · 30 days
11
4% of all-time downloads
All-time downloads
248
Public
Parameters
31.3B
18.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors18.4 GB · 100%
How the weights are stored.
U3230.7B · 98%
From the Hugging Face model README
MLX 4bit conversion of google/gemma-4-31b-it for Apple Silicon (mlx-vlm).
google/gemma-4-31b-it (license: gemma)mlx-vlm 0.6.3 — mlx_vlm.convert --hf-path google/gemma-4-31b-it --mlx-path . -q --q-bits 4 --q-group-size 64from mlx_vlm import load, generate
model, processor = load("ToPo-ToPo/gemma-4-31b-it-mlx-4bit")
This is a derivative of Google Gemma. Use is governed by the Gemma Terms of Use and the Gemma Prohibited Use Policy. Weights were converted/quantized to MLX format (modification notice per the Gemma Terms).
Use google/gemma-4-31b-it-assistant — Google's official MTP drafter for
this model. It loads directly in mlx-vlm (>= 0.6.3), needs no conversion, and speculative
decoding is lossless. Drafters are size-specific and not interchangeable across Gemma 4
variants.
chat_template.jinja differs from Google's Gemma 4 Canonical Chat Template
(2026-07-09) by one intentional change; everything else is untouched.
The canonical template suppresses the thinking channel at the start of a normal model
turn, but emits nothing after a tool response when enable_thinking is false. The model
may then open a thinking channel on its own, and a quantized model sometimes writes the
literal word thought into the answer. This patch gives the tool_response branch the
same suppression:
{%- elif ns.prev_message_type == 'tool_response' -%}
{%- if enable_thinking -%}
{{- '<|channel>thought\n' -}}
{%- else -%}
{{- '<|channel>thought\n<channel|>' -}}
{%- endif -%}
{%- endif -%}
Only that case changes — the other prompt paths render byte-identical to the canonical
template. To get stock behaviour, replace chat_template.jinja with the one from the
base model repo; the weights are unaffected.