Downloads · 30 days
19
8% of all-time downloads
mouri45/gemma-4-e2b-it-lite-mlx
gemma-4-e2b-it-lite-mlx is a text generation model from mouri45. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
Text-only, lightweight 4-bit MLX quantization of google/gemma-4-e2b-it, optimized for on-device inference on iPhone (8GB devices) and Apple Silicon Macs.
Downloads · 30 days
19
8% of all-time downloads
All-time downloads
229
Public
Parameters
4.6B
2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2 GB · 98%
How the weights are stored.
U324.6B · 100%
From the Hugging Face model README
Text-only, lightweight 4-bit MLX quantization of google/gemma-4-e2b-it, optimized for on-device inference on iPhone (8GB devices) and Apple Silicon Macs.
~2.05 GB — compared to 3.58 GB for the official mlx-community/gemma-4-e2b-it-4bit (which bundles the audio/vision towers in BF16).
mlx_lm.convert).audio_tower, vision_tower, embed_audio, embed_vision,
multi_modal_projector weights are not included. This checkpoint works with
text-only Gemma 4 runtimes (e.g. mlx-lm / mlx-swift-lm MLXLLM).embed_tokens_per_layer (per-layer embeddings, ~1.3 GB at 4-bit) which is
quantized to 2-bit / group size 64. The per-layer override is recorded in
config.json (quantization section).k_proj / v_proj / k_norm
(same layout as the official MLX conversion).Japanese smoke tests (greeting / Q&A / no repetition loops) show quality on par with the official 4-bit conversion on this recipe. A uniform 3-bit recipe of the same size was clearly worse and was rejected.
Gemma 4 is developed by Google DeepMind and released under the Apache License 2.0. This repository redistributes a quantized derivative under the same license.
pip install mlx-lm
mlx_lm.generate --model mouri45/gemma-4-e2b-it-lite-mlx --prompt "こんにちは!"
Created as part of the AppleSiliconLLM project (Issue0012).