Downloads · 30 days
33
42% of all-time downloads
thanglq150188/gemma4-e4b-mini
gemma4-e4b-mini is a text generation model from thanglq150188. Use it when you need the model to write or continue text. The card lists the license as gemma.
Vocabulary-pruned build of principled-intelligence/gemma-4-E4B-it-text-only: the token vocabulary is reduced to 69,246 entries covering English + Vietnamese, shrinking the embedding/LM-head so the q80 model fits comfo…
Downloads · 30 days
33
42% of all-time downloads
All-time downloads
78
Public
Repo size
5.2 GB
Likes
0
Public
Click a slice to open those files.
.gguf5.2 GB · 100%
From the Hugging Face model README
Vocabulary-pruned build of principled-intelligence/gemma-4-E4B-it-text-only: the token vocabulary is reduced to 69,246 entries covering English + Vietnamese, shrinking the embedding/LM-head so the q8_0 model fits comfortably on a 12 GB consumer GPU (~5.0 GB VRAM for weights, ~60 tok/s single-stream on an RTX 3060).
| Format | GGUF q8_0 (5.2 GB) |
| Architecture | Gemma-4 E4B text-only — 42 layers, hidden 2560, per-layer input embeddings, sliding+full hybrid attention |
| Vocab | 69,246 (EN+VI pruned) |
| Context | 131,072 tokens |
| Chat template | Gemma-4 `< |
pip install -U huggingface_hub (gives you the hf CLI). No Python libraries are needed to serve — llama.cpp is self-contained.hf download thanglq150188/gemma4-e4b-mini \
--local-dir ./models
# -> ./models/gemma4-e4b-envi-pruned-q8_0.gguf
# -> ./models/gemma4_chat.jinja
Uses the official llama.cpp server image with CUDA:
docker run -d --name gemma4-e4b \
--gpus all \
-p 8080:8080 \
-v "$(pwd)/models:/models" \
ghcr.io/ggml-org/llama.cpp:server-cuda \
-m /models/gemma4-e4b-envi-pruned-q8_0.gguf \
--host 0.0.0.0 --port 8080 \
-ngl 99 -c 131072 -np 16 \
-fa on --cache-reuse 256 \
--jinja --chat-template-file /models/gemma4_chat.jinja
Flag meanings, tune to taste:
-ngl 99 — offload all layers to GPU-c 131072 — context window (lower, e.g. -c 16384, to save VRAM)-np 16 — 16 parallel request slots (server throughput; each slot splits the context)-fa on — flash attention--cache-reuse 256 — prefix-cache reuse across requests (big win for shared system prompts)--jinja --chat-template-file ... — required: applies the Gemma-4 <|turn> chat template; without it the model receives a wrong prompt format and quality collapsesFor CPU-only, use the ghcr.io/ggml-org/llama.cpp:server image and drop --gpus all and -ngl 99.
# prebuilt releases: https://github.com/ggml-org/llama.cpp/releases (or build with cmake)
llama-server -m ./models/gemma4-e4b-envi-pruned-q8_0.gguf \
--host 0.0.0.0 --port 8080 -ngl 99 -c 131072 -np 16 \
-fa on --cache-reuse 256 --jinja \
--chat-template-file ./models/gemma4_chat.jinja
llama-server exposes an OpenAI-compatible endpoint at /v1:
curl http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
"messages": [
{"role": "user", "content": "Xin chào! Tóm tắt giúp tôi lợi ích của điện gió ngoài khơi."}
],
"max_tokens": 256, "temperature": 0.7
}'
Or from Python:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="none")
resp = client.chat.completions.create(
model="gemma4-e4b", # name is ignored by llama-server, any string works
messages=[{"role": "user", "content": "Việt Nam có bao nhiêu tỉnh thành?"}],
max_tokens=256,
)
print(resp.choices[0].message.content)
Health check: curl http://localhost:8080/health → {"status":"ok"} when the model is loaded.
gemma4_chat.jinja is mandatory; Gemma-3-style start_of_turn templates will not work with Gemma-4's <|turn> tokens.