Downloads · 30 days
14
25% of all-time downloads
yrrhall/Kimi-K2.5-MTP
Kimi-K2.5-MTP is a machine learning model from yrrhall. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Kimi K2.5 with Multi-Token Prediction (MTP) speculative decoding support for vLLM.
Downloads · 30 days
14
25% of all-time downloads
All-time downloads
57
Public
Parameters
1.1T
610 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors610 GB · 100%
How the weights are stored.
I321T · 96%
From the Hugging Face model README
Kimi K2.5 with Multi-Token Prediction (MTP) speculative decoding support for vLLM.
This model combines:
The MTP module enables speculative decoding where a lightweight prediction head drafts tokens that the main model can verify in parallel, improving throughput.
Requires modularml/vllm with MTP support for Kimi K2.5 (branch kimi-k25-mtp-780ba37).
vllm serve modularai/Kimi-K2.5-MTP \
--tensor-parallel-size 8 \
--trust-remote-code \
--speculative-config '{"model":"modularai/Kimi-K2.5-MTP","method":"mtp","num_speculative_tokens":1,"use_local_argmax_reduction":true}'
| Config | Output tok/s | TPOT p50 (ms) | Acceptance Rate |
|---|---|---|---|
| No speculation | 947 | 14.07 | — |
| MTP k=1 | 869 | 15.85 | ~39% |
Note: The MTP acceptance rate is low (~39%) because the MTP weights were not trained directly on this base model checkpoint. With properly matched MTP weights (trained via self-distillation on this exact checkpoint), acceptance rates of 80-90% are expected, yielding ~1.5-1.8x throughput improvement.
kimi_k25 (VLM wrapper around DeepSeek V3 architecture)enorm + hnorm (RMSNorm) → concat → eh_proj (Linear 2×7168 → 7168) → decoder layer (MLA attention + MoE) → shared_head (RMSNorm + LM head)This model uses weights from moonshotai/Kimi-K2.5 under the Modified MIT License and MTP weights from k-l-lambda/Kimi-K2.5-MTP under MIT.