Downloads · 30 days
31
15% of all-time downloads
slopops/KAT-Coder-V2.5-Dev-MTP-int4-AutoRound-SAR
KAT-Coder-V2.5-Dev-MTP-int4-AutoRound-SAR is a machine learning model from slopops. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
31
15% of all-time downloads
All-time downloads
205
Public
Parameters
1.5B
20.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors20.9 GB · 100%
How the weights are stored.
I324.3B · 71%
From the Hugging Face model README
This is a quantized, vLLM‑optimised version of KAT‑Coder‑V2.5‑Dev.
The model was post‑trained with int4 AutoRound (W4A16) and enhanced with a Multi‑Token Prediction (MTP) head from Qwen3.6‑35B‑A3B to enable speculative decoding in vLLM. It is designed for high‑throughput agentic coding workloads on hardware with 128 GiB unified memory (e.g. NVIDIA GB10).
[!IMPORTANT]
All benchmark scores reported in the KAT‑Coder‑V2.5 technical paper refer to the full BF16 model.
This int4 quantisation has not been re‑evaluated on standard benchmarks. It is expected to closely match the original quality, but individual results may vary.
| Property | Value |
|---|---|
| Base model | KAT‑Coder‑V2.5‑Dev (35 B total, 3 B active MoE) |
| Quantisation | int4 AutoRound (W4A16), group size 128 |
| MTP head | Added from Qwen3.6‑35B‑A3B (Apache 2.0) for speculative decoding |
| Calibration data | OpenCode Instruct (512 samples, sequence length 2048) |
| Quantisation tool | Spark Auto Round – an optimised fork of Intel® AutoRound |
| Deployment target | vLLM ≥ 0.19.0 with FlashInfer on GB10‑class hardware |
| Hugging Face Hub | slopops/KAT-Coder-V2.5-Dev-MTP-int4-AutoRound-SAR |
This model is intended for use with vLLM only.
We recommend using a recent official image (≥ 0.19.0) with FlashInfer support.
docker run --rm --gpus all --net=host --ipc=host \
vllm/vllm-openai:latest \
vllm serve slopops/KAT-Coder-V2.5-Dev-MTP-int4-AutoRound-SAR \
--port 8001 \
--host 0.0.0.0 \
--max-model-len 262144 \
--gpu-memory-utilization 0.55 \
--max-num-batched-tokens 16384 \
--max-num-seqs 8 \
--attention-backend flashinfer \
--enable-prefix-caching \
--enable-chunked-prefill \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--chat-template-kwargs '{"preserve_thinking":true}' \
--generation-config '{"temperature":0.6,"top_p":0.95,"top_k":-1,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
[!NOTE]
The server will be available athttp://localhost:8001/v1.
| Flag | Purpose |
|---|---|
--attention-backend flashinfer | Fast attention kernel, recommended for GB10 |
--speculative-config | Activates MTP with 3 speculative tokens |
--enable-auto-tool-choice + --tool-call-parser | Required for agentic tool‑use |
--chat-template-kwargs '{"preserve_thinking":true}' | Keeps thinking traces from previous turns |
--generation-config | Safe defaults for deterministic outputs |
--speculative-config flag.[!WARNING]
Performance claims (throughput, token savings, error reduction) refer to the original BF16 model as reported in the technical paper, or to our internal tests on a GB10 system. They are not guaranteed for every deployment. Always validate with your own workload.
If you use this quantised release, please cite the original KAT‑Coder work:
@misc{katcoder_v25_2026,
title={{KAT-Coder-V2.5 Technical Report}},
author={{KwaiKAT Team}},
year={2026},
month={July},
eprint={2607.05471},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/pdf/2607.05471}
}