Downloads · 30 days
11
8% of all-time downloads
hexoy/gemma-4-e2b-distilled
gemma-4-e2b-distilled is a image-text-to-text model from hexoy. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Distilled Gemma 4 E2B with all 35 language-model MLPs replaced by two-factor Monarch maps. The released BF16 model has 3,682,268,704 parameters, instead of 5,104,297,504 in the original Gemma 4 E2B.
Downloads · 30 days
11
8% of all-time downloads
All-time downloads
130
Public
Parameters
3.7B
7.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors7.4 GB · 100%
From the Hugging Face model README
Distilled Gemma 4 E2B with all 35 language-model MLPs replaced by two-factor
Monarch maps. The released BF16 model has 3,682,268,704 parameters, instead
of 5,104,297,504 in the original Gemma 4 E2B.
| Model | Status | Loaded weights | Serialized weights | Reduction vs dense BF16 |
|---|---|---|---|---|
| Dense Gemma 4 BF16 | Reference | 9.507 GiB | 9.543 GiB | - |
| Distilled Gemma 4 BF16 | Released | 6.859 GiB | 6.859 GiB | 27.86% |
| Distilled Gemma 4 + LoRA r8 BF16 | Experimental | 6.876 GiB | 6.877 GiB | 27.68% |
| Distilled Gemma 4 + INT8 linears | Released | 6.135 GiB | 6.136 GiB | 35.48% |
BF16 sizes use two bytes per parameter; serialized LoRA size was audited from the two safetensor shards. Values exclude activations and CUDA workspaces.
All rows used the same RTX PRO 6000, fixed batch size 32, seed 1234, 100
official examples, 10-shot prompts, and no chat template.
| Model | Status | GP-IRT accuracy | Raw accuracy | Runtime | Peak VRAM |
|---|---|---|---|---|---|
| Dense Gemma 4 BF16 | Reference | 39.23% | 29% | 28.50 s | 57.88 GiB |
| Distilled Gemma 4 BF16 | Released | 32.35% | 22% | 23.33 s | 55.23 GiB |
| Distilled Gemma 4 + LoRA r8 BF16 | Experimental | 33.25% | 23% | 24.93 s | 55.25 GiB |
| Distilled Gemma 4 + INT8 linears | Released | 30.57% | 21% | 23.14 s | 54.51 GiB |
Rank-8 LoRA improved the 35-layer source by 0.90 GP-IRT points and one correct
item, but missed the former +1.0 point or +2 item release gate. The
100-example benchmark is noisy.
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "hexoy/gemma-4-e2b-distilled"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
trust_remote_code=True,
dtype="auto",
device_map="auto",
)
The rank-8 result is published separately at
hexoy/gemma-4-e2b-monarch-35mlp-lora-r8.
It is a benchmark-recovery experiment: corrected 64-token generation repeated
phrases for text and remained malformed or inaccurate for images. It is not a
general text or multimodal recovery release.
f897353fca328b1cc5fd2e12d645773ca637f5f0ratmir-miftachov/gemma-distillationhexoy/gemma-4-e2b-monarch-35mlp-int8hexoy/gemma-4-e2b-monarch-35mlp-lora-r81435571b20dd26c073a535678975884154add5b8Detailed training and release evidence is retained privately.
Derived from google/gemma-4-E2B-it.
See NOTICE for the modification summary.