Downloads · 30 days
0
tokenfires/Qwen3.6-35B-A3B-Splash
Qwen3.6-35B-A3B-Splash is a text generation model from tokenfires. Use it when you need the model to write or continue text. It is set up for splash. The card lists the license as apache-2.0.
LM Studio per-model default (saved as ~/.lmstudio/.internal/user-concrete-model-default-config/incoai/Qwen3.6-35B-A3B-Splash.json, which LM Studio reads live):
Downloads · 30 days
0
Access
Public
Updated Sep 24, 2026
Repo size
20.9 GB
Likes
0
Public
Click a slice to open those files.
.bin20.9 GB · 100%
From the Hugging Face model README
| Setting | Use this | Why |
|---|---|---|
temperature | 0.6 | Qwen's thinking-mode value. Don't go greedy (0): Qwen's guidance is that it makes endless repetition more likely. |
top_p | 0.95 | Honored by the runtime. |
top_k | 20 | Honored. Must be an integer from 1 to 32. Anything higher, or "off", returns HTTP 400: Splash top-k must be an integer from 1 to 32. |
max_tokens | cap it, ~16k to 20k | This is your only loop guard. Without a cap, one runaway reasoning turn ran 112,000 tokens in about 6.5 minutes and never answered. |
repeat_penalty, presence_penalty, frequency_penalty | don't rely on them | The runtime ignores all three (byte-identical output with and without them). Setting them does nothing. |
| Thinking | on by default | /no_think doesn't stop it reasoning. The reasoning just goes to reasoning_content. |
LM Studio per-model default (saved as ~/.lmstudio/.internal/user-concrete-model-default-config/incoai/Qwen3.6-35B-A3B-Splash.json, which LM Studio reads live):
{
"preset": "",
"operation": {
"fields": [
{ "key": "llm.prediction.temperature", "value": 0.6 },
{ "key": "llm.prediction.topPSampling", "value": { "checked": true, "value": 0.95 } },
{ "key": "llm.prediction.topKSampling", "value": 20 },
{ "key": "llm.prediction.contextOverflowPolicy", "value": "rollingWindow" }
]
},
"load": { "fields": [] }
}
Every value above is honored except the ones in the "don't rely on them" row. Tested 2026-09-23 against LM Studio's Splash runtime.
A byte-for-byte mirror of incoai/Qwen3.6-35B-A3B-Splash, kept here as a pinned copy for my local agent stack. All credit for the packing, the DFlash 2 draft, and the Splash engine goes to Inco AI. Read their card first. This one only adds what I found running it.
Verified identical to upstream on 2026-09-23: the artifact_set_sha256 in manifest.json matches incoai's published manifest (0c465f96…2440a).
It's not a Transformers or MLX checkpoint. It only loads in Splash, or in LM Studio's Splash runtime.
| Path | Contents |
|---|---|
target/ | Qwen3.6-35B-A3B, 4-bit packed (splash-packed-q4-moe), 40 layers, 256 experts, 8 active per token |
draft/ | DFlash 2 draft model, 6 layers, proposes 7 tokens per step |
vision/ | Vision encoder |
tokenizer/ | Tokenizer and chat template |
manifest.json, layout.json | Artifact hashes, execution geometry, upstream revisions |
Upstream sources, from the manifest: target, tokenizer, and vision from mlx-community/Qwen3.6-35B-A3B-4bit @ 38740b8, draft from incoai/Qwen3.6-35B-A3B-DFlash2 @ 8e71350. About 20.9 GB on disk.
LM Studio: it shows up as qwen3.6-35b-a3b-splash and reports model format yuzu. This is how I run it.
Splash: upstream's command is splash serve --model incoai/Qwen3.6-35B-A3B-Splash. I haven't tried pointing Splash at this mirror's repo id, so use upstream's if you're going that route.
LM Studio on an M5 Max (128 GB), OpenAI-compatible endpoint, non-streaming: 1,000 tokens in 6.75 s, about 148 tok/s end to end with prefill included. Decode alone runs higher, around 170 tok/s on the same machine. For comparison, the Qwen3.8-27B MTPLX build I was running before did about 24 tok/s in the same agent.
I tested each parameter directly against the endpoint:
temperature, top_p, top_k. Also the per-model defaults in LM Studio's model config, which are picked up live.repeat_penalty, presence_penalty, frequency_penalty. Output at temperature 0 was byte-identical with and without them.So the usual repetition guard doesn't exist on this runtime (at least not yet). In my first evening of use it went into a reasoning loop: about 112k tokens of the same closing line, repeated inside the think block, and it never emitted an answer. At this speed that's six and a half minutes before anyone notices. If you run it behind an agent, cap max_tokens so a loop dies on its own. I'm running temperature 0.6, top_p 0.95, top_k 20, which is Qwen's thinking-mode recommendation. Don't go greedy, since Qwen's own guidance is that greedy decoding makes endless repetition more likely.
@lmstudio/sdk 1.5.0 hangs (it never resolves or rejects) on llm.listLoaded() whenever a Splash model is loaded, because its schema doesn't know the yuzu format. 2.0.0 fixes it. Upgrade before loading this.
Apache 2.0, same as upstream and the Qwen base model.