Downloads · 30 days
666
51% of all-time downloads
Anbeeld/Kimi-K2.7-Code-DSpark-GGUF
Kimi-K2.7-Code-DSpark-GGUF is a text generation model from Anbeeld. Use it when you need the model to write or continue text.
GGUF quantizations of novita DSpark draft model for Kimi-K2.7-Code.
Downloads · 30 days
666
51% of all-time downloads
All-time downloads
1.3K
Public
Repo size
19.1 GB
Likes
1
Public
Click a slice to open those files.
.gguf19.1 GB · 100%
From the Hugging Face model README
GGUF quantizations of novita DSpark draft model for Kimi-K2.7-Code.
Use with BeeLlama.cpp, a llama.cpp fork with advanced quantization features.
A DSpark speculator model for the Kimi-K2.7-Code base model, enabling faster
inference through speculative decoding. DSpark extends the DFlash parallel draft
backbone with two lightweight heads: a Markov logit-bias head (low-rank
intra-block token dependency) and a per-position confidence head (accept-rate
prediction). This checkpoint was trained in the Camelot-Ray online pipeline,
where the draft consumes hidden states streamed from a live Kimi-K2.7-Code vLLM
server.
This export is from Camelot exp38 checkpoint 3.
block_size=8mask_token_id=163608Online vLLM nightly spec-decode, greedy decoding, TP=8, Kimi-K2.7-Code verifier,
max_model_len=20000, cudagraphs enabled, and
fuse_allreduce_rms=false.
The table also includes Novita's public Eagle3-MLA draft
novita/kimi-k2.7-code-eagle3-mla under the same Kimi-K2.7-Code verifier,
TP=8, cudagraph, and fusion-off serving setup. Cells show
tok/s / speedup / accept_len. The standard rows use 6 prompts per benchmark;
code-extra rows use the full LiveCodeBench and SPEED-Bench coding manifests
with max_tokens=512.
| benchmark | rows | baseline tok/s | DSpark n=3 | DSpark n=7 | Novita Eagle3 n=3 | Novita Eagle3 n=7 | best |
|---|---|---|---|---|---|---|---|
| gsm8k | 6 | 132.0 | 282.2 / 2.14x / 2.937 | 309.1 / 2.34x / 3.659 | 281.7 / 2.13x / 2.941 | 277.3 / 2.10x / 3.595 | DSpark n=7 |
| math500 | 6 | 132.0 | 317.1 / 2.40x / 3.249 | 367.4 / 2.78x / 4.303 | 288.6 / 2.19x / 3.026 | 294.5 / 2.23x / 3.851 | DSpark n=7 |
| aime | 6 | 131.5 | 276.8 / 2.10x / 2.778 | 318.4 / 2.42x / 3.716 | 263.2 / 2.00x / 2.766 | 275.3 / 2.09x / 3.626 | DSpark n=7 |
| humaneval | 6 | 132.1 | 285.1 / 2.16x / 2.875 | 336.6 / 2.55x / 3.953 | 285.9 / 2.17x / 3.029 | 291.8 / 2.21x / 3.850 | DSpark n=7 |
| livecodebench | 121 | 129.8 | 227.5 / 1.75x / 2.306 | 231.0 / 1.78x / 2.696 | 219.5 / 1.69x / 2.342 | 198.5 / 1.52x / 2.593 | DSpark n=7 |
| speedbench_coding | 80 | 131.2 | 282.0 / 2.15x / 2.837 | 303.7 / 2.31x / 3.530 | 272.1 / 2.06x / 2.886 | 281.6 / 2.13x / 3.693 | DSpark n=7 |
Use DSpark with num_speculative_tokens=7 as the default for code, math, and
most reasoning traffic.
Requires a vLLM nightly with DSpark support:
uv pip install vllm --extra-index-url https://wheels.vllm.ai/nightly
vllm serve moonshotai/Kimi-K2.7-Code \
--tensor-parallel-size 8 \
--max-model-len 20000 \
--trust-remote-code \
--compilation-config='{"pass_config": {"fuse_allreduce_rms": false}}' \
--speculative-config '{
"model": "novita/kimi-k2.7-code-dspark",
"num_speculative_tokens": 7,
"method": "dspark"
}'
Known vLLM-nightly caveats, with workarounds:
scheduler_metadata must have shape (metadata_size) because the GPU-worker spec-decode path misses
fast_build=True when building draft attention metadata. Patch
vllm/v1/worker/gpu/spec_decode/speculator.py and
vllm/v1/worker/gpu/attn_utils.py to pass fast_build=True.--compilation-config='{"pass_config": {"fuse_allreduce_rms": false}}'.apply_verifier_norm=False, hidden_states = concat of aux
layers [1, 29, 57]