Downloads · 30 days
19
21% of all-time downloads
MichaelAnthony/gemma4-e2b-Snowfox-hf
gemma4-e2b-Snowfox-hf is a image-text-to-text model from MichaelAnthony. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This is the canonical merged BF16 Transformers checkpoint of SnowFox — a language-only LoRA merge built on Google's Gemma 4 E2B instruction QAT-derived model. Every SnowFox distribution (MLX FP16, MLX 4-bit, MLX 6-bit…
Downloads · 30 days
19
21% of all-time downloads
All-time downloads
92
Public
Parameters
5.1B
10.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors10.2 GB · 100%
From the Hugging Face model README
This is the canonical merged BF16 Transformers checkpoint of SnowFox — a language-only LoRA merge built on Google's Gemma 4 E2B instruction QAT-derived model. Every SnowFox distribution (MLX FP16, MLX 4-bit, MLX 6-bit, GGUF) is derived from this repository, so this is the package to use for full-precision Transformers inference or as the source for your own exports.
SnowFox is trained by Michael Anthony Falabella.
SnowFox is a language-only LoRA merge: the image and audio towers were frozen
during fine-tuning and are retained unchanged from the base model. Only the
language backbone received the SnowFox LoRA adaptation. The base is Google's
QAT-derived q4_0-unquantized checkpoint, which carries clipping parameters on
the multimodal towers that are preserved here.
| Property | Value |
|---|---|
| Total parameters | ~5.1B (with per-layer embeddings) |
| Effective parameters | ~2.3B |
| Weights format | BF16 |
| Checkpoint size | ~10.2 GB (model.safetensors) |
Note: Hugging Face's model page may report a smaller "params" figure for the quantized MLX derivatives of this model. That is a display artifact — those repos store weights as packed
uint32words (8× 4-bit / 5× 6-bit values per word) and HF counts each packed word as one parameter. The true count is unchanged (~5.1B total / ~2.3B effective).
google/gemma-4-E2B-it-qat-q4_0-unquantized6befbaca7398925921802abd1f277b495b78b738from transformers import AutoModelForCausalLM, AutoProcessor
model_id = "MichaelAnthony/gemma4-e2b-Snowfox-hf"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
| Package | Format | Notes |
|---|---|---|
gemma4-e2b-Snowfox-MLX | MLX FP16 | mlx-vlm ready |
gemma4-e2b-Snowfox-MLX-4bit | MLX 4-bit affine | ~3.55 GB |
gemma4-e2b-Snowfox-MLX-6bit | MLX 6-bit affine | ~4.71 GB |
gemma4-e2b-Snowfox-GGUF | GGUF | llama.cpp / Ollama |
Gemma 4 is Apache-2.0. This derivative package uses the Apache-2.0 license
declared by the pinned base model. See LICENSE and
NOTICE.md for the lineage and modification notice.