Downloads · 30 days
17
40% of all-time downloads
Wiself/Holodeck-Lounge-MTP
Holodeck-Lounge-MTP is a machine learning model from Wiself. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
TL;DR A working native Multi-Token Prediction (MTP) head on top of the Holodeck-Lounge merge. The upstream merge lineage shipped without a functional MTP head (the mtp.fc.weight tensor is absent after merge), so we re…
Downloads · 30 days
17
40% of all-time downloads
All-time downloads
42
Public
Parameters
9.7B
19.3 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors19.3 GB · 100%
From the Hugging Face model README
TL;DR A working native Multi-Token Prediction (MTP) head on top of the Holodeck-Lounge merge. The upstream merge lineage shipped without a functional MTP head (the
mtp.fc.weighttensor is absent after merge), so we restored it by sourcing only that single missing tensor —mtp.fc.weight— from Jackrong/Qwopus3.5-9B-Coder. The other 14 of the MTP head's 15 layers (themtp.layers.*blocks) were already present in the merge. We also patched in the froggeric/Qwen-Fixed-Chat-Templates fixed jinja template (v22.1), which improves reasoning structure and MTP acceptance in our evals. This model is the reference donor for our head fine-tune experiments.
mtp.fc.weight is missing, so native speculative
decoding is impossible from the source.mtp.layers.* transformer blocks); only the input-fusion projection
mtp.fc.weight was absent. We sourced that single tensor from
Jackrong/Qwopus3.5-9B-Coder, transplanted it into the Holodeck-Lounge
backbone, and verified it loads and runs end-to-end in llama.cpp
(draft-mtp) and vLLM (qwen3_5 MTP method). Backbone weights untouched.mtp.layers.* blocks come from the Holodeck-Lounge merge; only the
mtp.fc.weight fusion projection was sourced from
Jackrong/Qwopus3.5-9B-Coder (a model whose checkpoint carries the same
native Qwen3.5 MTP architecture).| Metric | Value |
|---|---|
| llama.cpp greedy / target-only acceptance | 59.0% |
| vLLM rejection-sampling acceptance | 58.55% |
| RS pos0 / pos1 / pos2 | 77.2% / 57.0% / 41.4% |
| Mean accepted length | 2.756 |
Eval: 30 prompts, temperature 0.7, max 256 tokens, concurrency 8. llama.cpp
numbers are greedy target-only acceptance (reasoning off); vLLM RS acceptance
from vllm:spec_decode_num_accepted_tokens / _num_draft_tokens. With
reasoning enabled per the Usage section, measured acceptance reaches the
60–69% range (n=30, same eval set).
Eval caveat: n=30 is a small sample — treat the 3-decimal precision as indicative; confidence intervals are wide.
With the recommended settings — draft-mtp in llama.cpp, qwen3_5 MTP in
vLLM — the restored head has consistently delivered a 60–69% draft-token
acceptance rate (65+% typical) in our testing. Real-world mileage will vary with workload
and hardware; treat 65% as a strong baseline, not a guarantee. The point of
this model isn't raw speed, but option: MTP when it pays off, plain greedy
decoding when it doesn't.
Where the win matters most:
The clean part: the host checkpoint already carried 14 of the MTP head's 15
layers, so this model ships at essentially the same size as its host — it was
always that big. The single-tensor transplant just completes the missing
mtp.fc.weight fusion layer, so you get the choice in one model, at no real
size cost.
Thanks to nightmedia for the Holodeck-Lounge backbone (which already carried 14
of the MTP head's 15 layers), to Jackrong/Qwopus3.5-9B-Coder for the
mtp.fc.weight tensor that completed the restore, and to froggeric for the
Qwen-Fixed-Chat-Templates v22.1 baked in here.
For best performance, use the baked-in froggeric template with reasoning enabled and preserved:
llama-server -m model.gguf \
--jinja --chat-template-file chat_template.jinja \
--reasoning-format deepseek --reasoning on --reasoning-preserve \
--spec-type draft-mtp --spec-draft-n-max 3
--spec-type draft-mtp --spec-draft-n-max 3--speculative-config '{"method":"mtp","num_speculative_tokens":3}'| Component | Source |
|---|---|
| Base architecture | Qwen/Qwen3.5-9B (Qwen team) |
| Backbone merge | nightmedia/Qwen3.5-9B-Holodeck-Lounge — 13-model creative-writing merge (DavidAU heretic family, armand0e, microsoft/Fara1.5-9B, Jackrong, et al.; see upstream card) |
MTP head — 14 mtp.layers.* blocks + mtp.norm/mtp.pre_fc_norm_* | Present in nightmedia/Qwen3.5-9B-Holodeck-Lounge merge |
MTP head — mtp.fc.weight (input-fusion projection, the missing 15th layer) | Jackrong/Qwopus3.5-9B-Coder |
| Chat template (v22.1) | froggeric/Qwen-Fixed-Chat-Templates |
License: apache-2.0. License chain (all components apache-2.0, verified): Qwen/Qwen3.5-9B → nightmedia merge → Jackrong MTP head → froggeric template.
--model-draft mtp-*.gguf).AI was used to draft this report.