Downloads · 30 days
46
10% of all-time downloads
majentik/Qwen3-Coder-Next-MLX-5bit
Qwen3-Coder-Next-MLX-5bit is a text generation model from majentik. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
5-bit weight-quantized MLX version of Qwen/Qwen3-Coder-Next, Qwen's 80B-A3B agentic coding MoE (512 experts, 10 active; hybrid Gated DeltaNet + Gated Attention; 256k context). Only ~3B parameters are active per token,…
Downloads · 30 days
46
10% of all-time downloads
All-time downloads
444
Public
Parameters
79.7B
54.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors54.8 GB · 100%
How the weights are stored.
U3279.7B · 100%
From the Hugging Face model README
5-bit weight-quantized MLX version of Qwen/Qwen3-Coder-Next,
Qwen's 80B-A3B agentic coding MoE (512 experts, 10 active; hybrid Gated DeltaNet +
Gated Attention; 256k context). Only ~3B parameters are active per token, so it runs
far faster than its 80B total suggests. Converted with mlx_lm — the canonical MLX
runtime for qwen3_next — and smoke-verified (chat + code probes) on Apple Silicon
with this exact payload before publishing. See PROVENANCE.md.
Approximate model size: ~55 GB
| Property | Value |
|---|---|
| Base Model | Qwen/Qwen3-Coder-Next |
| Parameters | 80 billion total (~3 billion active per token) |
| Architecture | MoE, hybrid Gated DeltaNet + Gated Attention (qwen3_next) |
| Modality | Text-only (code-focused) |
| Context Length | 256k tokens |
| License | Apache 2.0 |
| Weight Quantization | 5-bit affine, group size 64 (~55 GB) |
| Framework | MLX (Apple Silicon), mlx-lm >= 0.31 |
from mlx_lm import load, generate
model, tokenizer = load("majentik/Qwen3-Coder-Next-MLX-5bit")
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Write a Python function that merges two sorted lists."}],
add_generation_prompt=True, tokenize=False,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=512))
Or from the command line:
mlx_lm.generate --model majentik/Qwen3-Coder-Next-MLX-5bit --prompt "Refactor this function ..."
| Variant | Approx size | Use case |
|---|---|---|
| 2bit | ~25 GB | Smallest; quality floor |
| 3bit | ~35 GB | Low-RAM Macs |
| 4bit | ~45 GB | Balanced default |
| 5bit(https://huggingface.co/majentik/Qwen3-Coder-Next-MLX-5bit) | ~55 GB | Higher fidelity |
| 6bit | ~64 GB | Near-8bit quality |
| 8bit | ~84 GB | Reference fidelity |
Smoke verification covers load + short-form generation quality gates only; it is not a benchmark. For maximum fidelity use the largest variant that fits your unified memory (leave ~20% headroom for KV cache).