Downloads · 30 days
0
Hunterx/Kimi_K2.6_ToolCall_Template
Kimi_K2.6_ToolCall_Template is a machine learning model from Hunterx. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A community-fixed Jinja2 chat template for Kimi K2.6 (all quant formats — GGUF, MLX, etc.) that resolves tool calling failures across all major inference engines by replacing Kimi's native special tokens with a generi…
Downloads · 30 days
0
Access
Public
Updated May 8, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.jinja11.8 KB · 58%
From the Hugging Face model README
A community-fixed Jinja2 chat template for Kimi K2.6 (all quant formats — GGUF, MLX, etc.) that resolves tool calling failures across all major inference engines by replacing Kimi's native special tokens with a generic format that llama.cpp, ik_llama.cpp, oMLX, and LM Studio can actually parse.
Drop-in replacement for the default tokenizer_config.json chat template.
Kimi K2.6 uses unique special tokens for tool calls (<|tool_call_begin|>, <|tool_call_argument_begin|>, <|tool_calls_section_begin|>, etc.). The only inference engine with a native parser for these tokens is vLLM. Every other engine — llama.cpp, ik_llama.cpp, oMLX, LM Studio, KoboldCpp — either strips these tokens silently or fails to parse them, resulting in:
Native tool parser failed (ValueError: No tool call found.) errors in oMLX<tool_call> JSON format that every inference engine's generic parser understands.| Issue | Stock / v1 Behavior | v2 Fix |
|---|---|---|
| Native special tokens for tool calls | Outputs <|tool_call_begin|>, <|tool_call_argument_begin|>, etc. — only parseable by vLLM | Outputs generic <tool_call>{"name": "...", "arguments": {...}}</tool_call> format understood by all engines |
| Tool response format | Uses ## Return of {{ tool_call_id }} — non-standard | Uses <tool_response>...</tool_response> tags matching generic format |
| System prompt tool instructions | Shows native token format as example | Shows <tool_call> JSON format so model follows the generic pattern |
| Missing function name | Stock template never outputs function.name in tool calls | Function name included in JSON output |
| Fix | Description |
|---|---|
Auto-close <think> before tool calls | Prevents reasoning from leaking into structured tool output |
| Strict tool calling rules | System prompt includes behavioral rules for reliable tool use |
| String-form argument handling | Handles model outputting arguments as JSON string instead of object |
Both </think> and </thinking> recognized | Supports both close tag variants |
| Think toggles | <|think_on|> / <|think_off|> in any message |
| Developer role | Supports developer role messages |
| Cross-runtime compatible | No .get() calls, no is sequence — works on limited Jinja runtimes |
| Historical reasoning hidden | Previous turns' think blocks stripped to save tokens |
<|im_user|>, <|im_assistant|>, <|im_system|>, <|im_middle|>, <|im_end|><|media_begin|>, <|media_content|>, <|media_pad|>, <|media_end|><|kimi_k25_video_placeholder|>preserve_thinking flag for debuggingtools_ts_str support for pre-formatted tool stringsmessage.nameThis template changes the tool call output format from Kimi's native tokens to a generic format. The model was trained with native tokens, so there is a possibility it may occasionally revert to its trained format. In testing, providing clear format instructions in the system prompt (which this template does) is sufficient to guide the model to use the generic format consistently. If you see native tokens leaking through, try lowering temperature or adding reinforcing instructions to your system prompt.
llama-server -m your-kimi-model.gguf \
--jinja \
--chat-template-file kimi_k2.6_fixed_template_v2.jinja
ik_llama_server -m your-kimi-model.gguf \
--jinja \
--chat-template-file kimi_k2.6_fixed_template_v2.jinja
For thinking control via API:
ik_llama_server -m your-kimi-model.gguf \
--jinja \
--chat-template-file kimi_k2.6_fixed_template_v2.jinja \
--chat-template-kwargs '{"enable_thinking":false}' \
--reasoning-budget 0
Copy the contents of kimi_k2.6_fixed_template_v2.jinja into the chat_template field of your model's tokenizer_config.json.
Go to My Models → model settings → Prompt Template and paste the template contents.
Stock template output (broken on most engines):
<|tool_calls_section_begin|>
<|tool_call_begin|>read
call_abc123<|tool_call_argument_begin|>{"filePath": "/src/main.py"}<|tool_call_end|>
<|tool_calls_section_end|>
v2 template output (works everywhere):
<tool_call>
{"name": "read", "arguments": {"filePath": "/src/main.py"}}
</tool_call>
tools_ts_str pre-formatted tool string path is preserved but untested with the generic format.Template fixes by Hunterx, based on the community fix pattern established by:
Same license as the original Kimi K2.6 model.