Downloads · 30 days
852
32% of all-time downloads
AtomicChat/Qwen3.5-9B-MLX-4bit
Qwen3.5-9B-MLX-4bit is a text generation model from AtomicChat. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
Downloads · 30 days
852
32% of all-time downloads
All-time downloads
2.7K
Public
Parameters
9B
5.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5 GB · 100%
How the weights are stored.
U329B · 100%
From the Hugging Face model README
Qwen3.5 9B, self-quantized to MLX by Atomic Chat. Built straight from Qwen's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
[!NOTE] These MLXs are self-quantized from the original weights, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
| Property | Value |
|---|---|
| Base model | Qwen/Qwen3.5-9B |
| Parameters | 9.7B |
| Layers | 32 |
| Context length | 262,144 tokens (256K) |
| Vocabulary | 248,320 |
| Modalities | Text, Image |
| Architecture | Dense decoder, 16 attention heads over 4 KV heads, Qwen3_5ForConditionalGeneration |
| This repo | MLX weights |
AtomicChat/Qwen3.5-9B-MLX-4bit and hit Use this model.mlx_lm.generate --model AtomicChat/Qwen3.5-9B-MLX-4bit --prompt "Hello" --max-tokens 512mlx_lm.server --model AtomicChat/Qwen3.5-9B-MLX-4bit --port 8080| Parameter | Value |
|---|---|
| temperature | 1.0 |
| top_p | 0.95 |
| top_k | 20 |
| min_p | 0.0 |
| repetition_penalty | 1.0 |
Qwen's recommended sampling configuration for Qwen/Qwen3.5-9B.
Qwen/Qwen3.5-9B (original weights).mlx_lm.convert on our pipeline.Original model by Qwen, released under the Apache 2.0 license. Full terms: Apache 2.0. Quantized by Atomic Chat.