Downloads · 30 days
13
8% of all-time downloads
LLMWildling/gemma-4-100b-a10b-coder
gemma-4-100b-a10b-coder is a text generation model from LLMWildling. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as gemma.
gemma-4-100b-a10b-coder is a Gemma 4 based coder model for software engineering, code editing, Q/A, tool use, and long-context assistant workflows.
Downloads · 30 days
13
8% of all-time downloads
All-time downloads
160
Public
Parameters
59.3B
62.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors62.3 GB · 100%
How the weights are stored.
U856.3B · 95%
From the Hugging Face model README
gemma-4-100b-a10b-coder is a Gemma 4 based coder model for software engineering, code editing, Q/A, tool use, and long-context assistant workflows.
gemma-4coder100.3B10.0B38200000 tokens in the listed vLLM configurationCUDA_VISIBLE_DEVICES=0,1 vllm serve /path/to/gemma-4-100b-a10b-coder \
--served-model-name vllm/doobee \
--host 0.0.0.0 \
--port 23333 \
--dtype bfloat16 \
--tensor-parallel-size 2 \
--enable-expert-parallel \
--max-model-len 200000 \
--gpu-memory-utilization 0.96 \
--trust-remote-code \
--reasoning-parser gemma4 \
--tool-call-parser gemma4 \
--enable-auto-tool-choice \
--default-chat-template-kwargs '{"enable_thinking": true}' \
--generation-config vllm \
--language-model-only \
--skip-mm-profiling \
--max-num-seqs 1 \
--max-num-batched-tokens 8192
For clients that should not receive reasoning text, send
"include_reasoning": false in chat-completion requests.
config.jsongeneration_config.jsontokenizer.jsontokenizer_config.jsonchat_template.jinjamodel.safetensors.index.jsonThis model is released under the Gemma license.