Downloads · 30 days
0
QmogAI/gemma4-e2b-int8
gemma4-e2b-int8 is a text generation model from QmogAI. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This is the ready-to-run model file for gemma4.c, a small educational implementation of Gemma 4 E2B inference on CPU.
Downloads · 30 days
0
Access
Public
Updated Aug 26, 2026
Repo size
5 GB
Likes
1
Public
Click a slice to open those files.
.bin5 GB · 100%
From the Hugging Face model README
This is the ready-to-run model file for gemma4.c, a small educational implementation of Gemma 4 E2B inference on CPU.
The project's exporter.py converts the original Google Gemma 4 E2B QAT checkpoint into one 4.7 GiB file containing the model configuration, tokenizer, and weights. During export, matrix weights are quantized to int8 with FP16 scales.
hf download QmogAI/gemma4-e2b-int8 gemma4-E2B-int8.bin --local-dir .
git clone https://github.com/ryanssenn/gemma4.c
cd gemma4.c
make
./run -m ./gemma4-E2B-int8.bin -t 1.0 -n 256 "Why is the sky blue?"
If you would rather create the file from the original checkpoint:
python3 -m pip install -r requirements.txt
python3 exporter.py /path/to/gemma-4-E2B-it-qat-q4_0-unquantized -o ./gemma4-E2B-int8.bin
The runtime is made for x86-64 CPUs with AVX2 and FMA. AVX-512 VNNI is used when available. It is intended for learning and experimentation.