Downloads · 30 days
1M
43% of all-time downloads
mudler/KAT-Coder-V2.5-Dev-APEX-GGUF
KAT-Coder-V2.5-Dev-APEX-GGUF is a machine learning model from mudler. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<div style="background-color: f59e0b; color: white; padding: 20px; border-radius: 10px; text-align: center; margin: 20px 0;" <h2 style="color: white; margin: 0 0 10px 0;"⚡ Each donation = another big MoE quantized</h2…
Downloads · 30 days
1M
43% of all-time downloads
All-time downloads
2.4M
Public
Repo size
212 GB
Likes
62
Public
Click a slice to open those files.
.gguf143 GB · 100%
From the Hugging Face model README
APEX (Adaptive Precision for EXpert Models) quantizations of Kwaipilot/KAT-Coder-V2.5-Dev — Kwaipilot's Mixture-of-Experts model for agentic coding.
Brought to you by the LocalAI team | APEX Project | Technical Report
| File | Profile | Best For |
|---|---|---|
| KAT-Coder-V2.5-Dev-APEX-I-Balanced.gguf | I-Balanced | Best overall — imatrix-enhanced |
| KAT-Coder-V2.5-Dev-APEX-I-Quality.gguf | I-Quality | Highest quality with imatrix |
| KAT-Coder-V2.5-Dev-APEX-Quality.gguf | Quality | Highest quality (no imatrix) |
| KAT-Coder-V2.5-Dev-APEX-Balanced.gguf | Balanced | General purpose |
| KAT-Coder-V2.5-Dev-APEX-I-Compact.gguf | I-Compact | Consumer GPUs, imatrix-enhanced |
| KAT-Coder-V2.5-Dev-APEX-Compact.gguf | Compact | Consumer GPUs |
| KAT-Coder-V2.5-Dev-APEX-I-Mini.gguf | I-Mini | Smallest viable, fastest inference |
APEX is a quantization strategy for Mixture-of-Experts (MoE) models. It classifies tensors by role (routed expert, shared expert, attention) and applies a layer-wise precision gradient — edge layers (first/last 5) get higher precision, middle layers compress more aggressively. I-variants use diverse imatrix calibration (chat, code, reasoning, tool-calling, agentic traces, Wikipedia).
In MoE models the routed-expert FFN tensors dominate the weight budget but only ~8/256 experts activate per token, so APEX compresses middle-layer experts hardest while preserving edge layers, attention, and the always-active shared expert.
See the APEX project for full details.
Note: the config advertises an image token, but the released checkpoint ships no vision encoder weights, so these are text-only GGUFs (no mmproj).
local-ai run mudler/KAT-Coder-V2.5-Dev-APEX-GGUF@KAT-Coder-V2.5-Dev-APEX-I-Balanced.gguf
APEX is brought to you by the LocalAI team. Built on llama.cpp. Base model by Kwaipilot.