Downloads · 30 days
7.8K
53% of all-time downloads
mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit
KAT-Coder-V2.5-Dev-OptiQ-4bit is a text generation model from mlx-community. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). Try the Lab · All OptiQ quants · Docs
Downloads · 30 days
7.8K
53% of all-time downloads
All-time downloads
14.6K
Public
Parameters
34.7B
22 GB on disk
Likes
20
Trending 1
Click a slice to open those files.
.safetensors21.9 GB · 100%
How the weights are stored.
U3234.7B · 100%
From the Hugging Face model README
Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). Try the Lab · All OptiQ quants · Docs
An OptiQ mixed-precision MLX quant of
KAT-Coder-V2.5-Dev, a qwen3_5_moe coding model (256-routed-expert sparse
MoE with hybrid linear + full attention).
static build — per-layer bit-widths assigned to a 4.5
target bits-per-weight: 400 projections at 4-bit, 111 at 8-bit.pip install -U optiq
qwen3_5_moe loads under stock mlx-lm too, but optiq serve adds mixed-
precision loading, KV-cache quantization, and the OptiQ Lab.
optiq serve --model mlx-community/KAT-Coder-V2.5-Dev-OptiQ-4bit
Then use the OpenAI-compatible endpoint at http://localhost:8000/v1, the
OptiQ Lab, or point optiq code at it. This
is a reasoning coder — it thinks before it answers.