Downloads · 30 days
36
25% of all-time downloads
DeepDavid/gemma-4-e4b-it-qcache-q8
gemma-4-e4b-it-qcache-q8 is a machine learning model from DeepDavid. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as gemma.
Quantized deployment artifacts for google/gemma-4-e4b-it.
Downloads · 30 days
36
25% of all-time downloads
All-time downloads
145
Public
Repo size
12.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors7.7 GB · 61%
From the Hugging Face model README
Quantized deployment artifacts for google/gemma-4-e4b-it.
qcache-q8_0.v1.gguf — GGML Q8_0 tensors for the model's linear projections,
stored in a GGUF container keyed by the original checkpoint tensor paths
(no metadata KVs). Produced by in-memory quantization of the bf16
checkpoint with a candle-based
loader.aux-tensors.safetensors — everything the cache does not carry:
embeddings, norms and other non-projection tensors, in their original
dtypes.Format note: this is not a llama.cpp-compatible GGUF — tensors keep their original checkpoint names and only linear projections are quantized. Load it with a runtime that pairs the cache with the auxiliary safetensors.
Weights are redistributed under the same terms as the base model.