Downloads · 30 days
21
12% of all-time downloads
LLMWildling/gemma4-52b-a6b
gemma4-52b-a6b is a text generation model from LLMWildling. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as gemma.
gemma4-52b-a6b is a production MoE expansion of unsloth/Gemma-4-26B-A4B-it. It was built with a two-GPU local training setup and expanded with additional specialist capacity for software engineering, repository reason…
Downloads · 30 days
21
12% of all-time downloads
All-time downloads
178
Public
Parameters
28B
29.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors29.3 GB · 100%
How the weights are stored.
U823.7B · 85%
From the Hugging Face model README
gemma4-52b-a6b is a production MoE expansion of
unsloth/Gemma-4-26B-A4B-it. It was built with a two-GPU local training setup
and expanded with additional specialist capacity for software engineering,
repository reasoning, agentic workflows, and general instruction following.
This is a full model checkpoint, not a LoRA adapter.
unsloth/Gemma-4-26B-A4B-it50.6B6B128 base experts + 128 added specialist experts256k tokens131k tokens0.5 to 0.7Use the included Gemma4 chat template. Thinking mode is intended to be enabled.
CUDA_VISIBLE_DEVICES=0 \
vllm serve . \
--served-model-name gemma4-52b-a6b \
--host 0.0.0.0 \
--port 23333 \
--dtype bfloat16 \
--tensor-parallel-size 1 \
--max-model-len 131072 \
--gpu-memory-utilization 0.88 \
--trust-remote-code \
--reasoning-parser gemma4 \
--tool-call-parser gemma4 \
--enable-auto-tool-choice \
--chat-template ./chat_template.jinja \
--default-chat-template-kwargs '{"enable_thinking": true}'
OpenAI-compatible endpoint:
http://localhost:23333/v1
This release is intended for production evaluation of the expanded Gemma4 MoE line. Feedback is welcome through the model discussion page or the usual Hugging Face repository feedback channels.