Downloads · 30 days
692
66% of all-time downloads
logic65/whittle-next
whittle-next is a machine learning model from logic65. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<p align="center"<img src="whittle.svg" width="640" alt="Whittle"</p
Downloads · 30 days
692
66% of all-time downloads
All-time downloads
1.1K
Public
Parameters
16.4B
160 GB on disk
Likes
2
Public
Click a slice to open those files.
.gguf89.7 GB · 56%
From the Hugging Face model README
☕ Support this work
Whittle is built by one person on a grocery budget and rented GPU hours. If this research is useful to you, or you want to see it finished: ko-fi.com/davida81328. Every hour of GPU time goes straight into the next checkpoint, and every checkpoint, table and log lands in these repos.
⚠️ Research artifact. A 19.8B qwen4exp-architecture model built by weight surgery on Whittle-tri-14.7B (Qwen3.8-27B depth-compressed 64 -> 32 layers by parallel-compose merging, FFN width uncut, Apache-2.0), then repaired by SFT. It now holds a conversation, follows the chat template, writes fenced code, and stops cleanly — but it is factually thin and its arithmetic is approximate. Treat it as an architecture demonstrator, not an assistant.
Earlier Whittle-Next release (research artefact). The Whittle-Next line continued with
Whittle-Next-27B-A3B, parent of the current
Whittle-Qwen-3.8-35B-A3B. This repo holds the GGUF builds of the 19.8B
lineage; the router104 file below is the build the
Qwen3.8-Whittle-Next-19.8B-A11B-chat card points to.
Serve every file with the settings under Run it; they are required, not suggestions.
| file | what it is | recommended |
|---|---|---|
whittle-next-qwen4exp-sft-PLE4B-Q4_K_M.gguf | SFT + woken hyper-connections + trained shared-expert gates + 4B n-gram memory | ✅ yes |
whittle-next-qwen4exp-router104-PLE4B-Q4_K_M.gguf | the above plus jointly-trained routers at k=104 — better offline metrics, worse behaviour (see Measured) | experimental |
whittle-next-qwen4exp-HC-Q4_K_M.gguf | hyper-connections only, no n-gram memory | ablation |
whittle-next-qwen4exp-HC-PLE4B-f16.gguf | f16, n-gram memory, pre-SFT | ablation |
Also in this repo, not described on this card (names as in the file listing; sizes as the Hub reports them, binary GiB):
whittle-next-qwen4exp-SFT-PLE4B-Q4_K_M.gguf (capital SFT, 12.31 GiB) — a second SFT build next to the recommended lowercase sft file (12.29 GiB); the two files differ in size and this card does not say how they differ.model-00001.safetensors … model-00008.safetensors, model.safetensors.index.json) with config.json, generation_config.json, chat_template.jinja, tokenizer.json, tokenizer_config.json.hc_init.pt, mini_next.py.next-v2/ (hc_ple.pt, ngram_table_4B.pt, train_phase1.log) and next-v3/ (hc_ple_attn.pt, ngram_table_4B.pt, train_final.log).recovery/phaseA_step1250/ and recovery/phaseC_step100/ (hc.pt, ple.pt, routers.pt).| build | mode | clean stops | looping answers |
|---|---|---|---|
| SFT (recommended) | thinking off | 5/6 | 1/6 |
| SFT (recommended) | thinking on | 4/6 | 1/6 |
| router104 | thinking off | 4/6 | 2/6 |
| router104 | thinking on | 2/6 | — over-thinks, ran out of budget |
With the recommended sampling the remaining loop disappears: longform, explanation, code and list
probes all returned finish_reason=stop with 4-gram repetition 0.000 (one short story at 0.38).
Why router104 is not the default, despite better numbers. Training the routers jointly with the shared-expert gates, hyper-connections and n-gram projections — and at the k they serve — produced the best offline metrics this project has recorded (held-out CE 4.1466 → 3.9745, fact battery 4/5 → 5/5). But served, it over-thinks and repeats more. The training-time gate was selecting on cross-entropy and a short greedy battery, neither of which measures paragraph-length generation; repetition on that gate rose 0.057 → 0.093 over the same window while CE improved. The router result is real and reproducible — it is a training-objective lesson, not a serving win.
Mechanically conversational; not yet substantively reliable. It takes a turn, answers, and stops — and the content underneath is often wrong. Verified single-turn probes (recommended build, k=104, serving sampler): Paris ✅, a complete valid fenced HTML page ✅, a coherent non-repeating paragraph ✅ — against "the sky is blue because sunlight shines through the clouds" ❌, 17+25 = 32 ❌, and "list exactly 5 fruits" sometimes answered "1, 2, 3, 4, and 5 are fruits" ❌ (it hears the format and misses the substance).
Untested: every probe is single-turn. Multi-turn context retention — arguably the real test of "conversational" — has not been measured, and we make no claim about it.
The failure mode has moved from broken generation to a small model with damaged knowledge.
19.775B parameters, 32 layers × 5120, 3:1 GDN:full-attention, 240 experts (k=104 recommended,
58 trained), 4 hyper-connection residual streams, per-layer n-gram memory over a 6.25M-row × 640
table (≈4B parameters, host-offloadable). Requires a llama.cpp with qwen4exp support; the GGUFs
declare output_gate_type: silu, which transformers' qwen4_exp config now supports natively —
the Qwen3.5-derived GDN weights need a SiLU output gate, not the sigmoid a Flash-Next model uses.
Serving settings — these are REQUIRED, not suggestions:
llama-server -m whittle-next-qwen4exp-sft-PLE4B-Q4_K_M.gguf -ngl 99 -c 8192 --jinja \
--override-kv qwen4exp.expert_used_count=int:104
Request body: temperature 0.7, top_p 0.8, top_k 20, repeat_penalty 1.05.
Two settings do almost all the work, and both were measured on this build:
temperature 0 a paragraph-length answer degenerates
(4-gram repetition 0.885 — "the ocean is a combination of water and water…").
At the settings above the same prompt scores 0.000 and ends with finish_reason=stop.
Greedy decoding is the single largest cause of looping in this model.--override-kv above). Raising k from the trained 58 to 104 is a
config-only change that fixed list termination, restored task engagement (a "build a page" request
went from a fabricated URL to real fenced HTML), and removed intra-list repetition — with zero
gradient steps.Reasoning is optional: pass chat_template_kwargs: {"enable_thinking": false} for short factual
turns. With thinking on, allow ≥700 tokens — the think block is verbose.
Lineage: Qwen3.8-27B → Whittle-tri-14.7B (depth-compressed 64 -> 32 layers by parallel-compose merging, FFN width uncut) → this qwen4exp build, repaired by SFT. Apache-2.0.
David Aylward (logic65) & Claude (Anthropic) — designed, debugged and trained together.