Downloads · 30 days
238
20% of all-time downloads
AtomicChat/Qwen3-Coder-Next-DFlash-GGUF
Qwen3-Coder-Next-DFlash-GGUF is a text generation model from AtomicChat. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as mit.
Downloads · 30 days
238
20% of all-time downloads
All-time downloads
1.2K
Public
Repo size
1.5 GB
Likes
0
Public
Click a slice to open those files.
.gguf510 MB · 100%
From the Hugging Face model README
Qwen3 Coder Next Dflash, self-quantized to GGUF by Atomic Chat. Built straight from Z Lab's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
[!NOTE] These GGUFs are self-quantized from the original weights, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
[!IMPORTANT] Always pass
--jinjaso the Qwen3 Coder Next Dflash chat template is applied. Without it the model can emit malformed turns.
| Property | Value |
|---|---|
| Base model | z-lab/Qwen3-Coder-Next-DFlash |
| Parameters | 0.5B |
| Layers | 8 |
| Context length | 262,144 tokens (256K) |
| Vocabulary | 151,936 |
| Modalities | Text |
| Architecture | Dense decoder, 32 attention heads over 4 KV heads, DFlashDraftModel |
| This repo | GGUF quants (imatrix). Quants: Q8_0 |
| Quant | Size | Notes |
|---|---|---|
Q8_0 | 0.5 GB | Effectively lossless, reference quality. |
[!TIP] Pick the largest file that fits your (V)RAM with room for context.
Q8_0is the sweet spot for most setups;Q6_KorQ8_0for maximum fidelity.
Run Qwen3 Coder Next Dflash locally with:
AtomicChat/Qwen3-Coder-Next-DFlash-GGUF, pick a quant, hit Use this model.llama-server -hf AtomicChat/Qwen3-Coder-Next-DFlash-GGUF:Q8_0 --jinja -c 8192ollama run hf.co/AtomicChat/Qwen3-Coder-Next-DFlash-GGUF:Q8_0| Parameter | Value |
|---|---|
| temperature | 0.0 |
Z Lab's recommended sampling configuration for z-lab/Qwen3-Coder-Next-DFlash.
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server
./llama.cpp/build/bin/llama-server \
-hf AtomicChat/Qwen3-Coder-Next-DFlash-GGUF:Q8_0 \
--jinja -ngl 99 -c 8192 -fa on
z-lab/Qwen3-Coder-Next-DFlash (original weights).--imatrix.Original model by Z Lab, released under the MIT license. Quantized by Atomic Chat.