Downloads · 30 days
302
20% of all-time downloads
systemsofrecord/Qwen3-Coder-30B-A3B-Instruct-W4A16
Qwen3-Coder-30B-A3B-Instruct-W4A16 is a text generation model from systemsofrecord. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
W4A16 (INT4 group-128 weights + FP16 activations) quantization of Qwen/Qwen3-Coder-30B-A3B-Instruct.
Downloads · 30 days
302
20% of all-time downloads
All-time downloads
1.5K
Public
Parameters
30.8B
16.7 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.7 GB · 100%
How the weights are stored.
I3229.9B · 97%
From the Hugging Face model README
W4A16 (INT4 group-128 weights + FP16 activations) quantization of Qwen/Qwen3-Coder-30B-A3B-Instruct.
pack-quantized
format.The point of this build is to fit this 30B-A3B MoE onto small, FP8-less GPUs like the A2, where BF16 (~57 GB) and INT8 (~30 GB) don't fit. At 4-bit the checkpoint is ~16 GB (~8 GB/GPU at TP=2), running via the Marlin INT4 kernel — which, unlike FP8 / W4AFP8, works on Ampere.
| Quantized → INT4 (g128, symmetric) | Kept in BF16 |
|---|---|
| all 128 routed experts × 48 layers | token embeddings, lm_head |
attention q/k/v/o projections | MoE router gates, all norms |
Only transformer Linear weights are quantized; the embedding, output head, router gates,
and norms stay BF16 for quality. It remains a standard Qwen3MoeForCausalLM — full GQA
attention, SwiGLU, 128 experts / 8 active — so it uses vLLM's mainstream MoE path.
vllm serve systemsofrecord/Qwen3-Coder-30B-A3B-Instruct-W4A16 \
--tensor-parallel-size 2 \
--dtype float16 \
--max-model-len 32768 \
--gpu-memory-utilization 0.90 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
Qwen3MoeForCausalLM support (≥ 0.25).Verified: loaded and generated correct code on 2× NVIDIA A2 under vLLM 0.25.1 — ~7.9 GB weights/GPU at TP=2, Marlin wNa16 MoE kernel, CUDA graphs captured cleanly.
W4A16 — weights 4-bit int, group_size=128, symmetric; activations unquantized.pack-quantized (INT4 packed into INT32 + group scales).lm_head, embed_tokens, MoE router gates, norms.Apache-2.0, inherited from the base model Qwen/Qwen3-Coder-30B-A3B-Instruct. This repository only redistributes a quantized copy of those weights; all model capabilities and credit belong to the Qwen team.