Downloads · 30 days
22
6% of all-time downloads
ltpla/gemma-4-e4b-it-noaudio-4bit
gemma-4-e4b-it-noaudio-4bit is a any-to-any model from ltpla. Use it for the any-to-any task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as gemma.
This is a modified version of mlx-community/gemma-4-e4b-it-4bit. The Universal Speech Model encoder (audiotower.) and audio embedder (embedaudio.) — 754 weight keys, ~270 MB at 4-bit — have been removed; config.json h…
Downloads · 30 days
22
6% of all-time downloads
All-time downloads
375
Public
Parameters
7.7B
4.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.6 GB · 99%
How the weights are stored.
U327.5B · 97%
From the Hugging Face model README
This is a modified version of mlx-community/gemma-4-e4b-it-4bit. The Universal Speech Model encoder (audio_tower.*) and audio embedder (embed_audio.*) — 754 weight keys, ~270 MB at 4-bit — have been removed; config.json has audio_config and audio_token_id dropped and has_audio set to false. The text model and vision tower are unchanged.
Useful when audio input is not needed and disk/memory footprint matters (e.g. on systems with 16 GB unified memory). Audio prompts will fail at the model level — the audio tower is gone. Text-only and image inputs work exactly as the original.
Gemma is provided under and subject to Google's Gemma Terms of Use and Gemma Prohibited Use Policy. By using, modifying, or distributing this model you agree to those terms, including the prohibited-use restrictions. This work is a modification; the original Gemma 4 model card is at google/gemma-4-e4b-it.
audio_tower.* and embed_audio.* weightsaudio_config and audio_token_id from config.jsonhas_audio: falsemodel.safetensors.index.jsonpip install -U mlx-vlm
python -m mlx_vlm.generate \
--model ltpla/gemma-4-e4b-it-noaudio-4bit \
--max-tokens 100 --temperature 0.0 \
--prompt "Describe this image." --image <path_to_image>