Downloads · 30 days
537
12% of all-time downloads
Achilles1089/fable-coder-35B-A3B
fable-coder-35B-A3B is a text generation model from Achilles1089. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A sovereign, open-weights agentic coding model by Dappit Labs. 35B Mixture-of-Experts (≈3B active), built by layering Claude Fable-5 / Opus-4.8 agentic tool-use behavior onto an abliterated, Opus-4.7-reasoning-distill…
Downloads · 30 days
537
12% of all-time downloads
All-time downloads
4.6K
Public
Parameters
36B
71.9 GB on disk
Likes
14
Public
Click a slice to open those files.
.safetensors71.9 GB · 100%
From the Hugging Face model README
A sovereign, open-weights agentic coding model by Dappit Labs. 35B Mixture-of-Experts (≈3B active), built by layering Claude Fable-5 / Opus-4.8 agentic tool-use behavior onto an abliterated, Opus-4.7-reasoning-distilled Qwen3.6-35B-A3B.
Built by Dappit Labs · @dappitdotio Trained on hardware provided by Manifest Network. 🙏
⚠️ Numbers are from our own harness (see Evaluation); nothing here is a claim against official leaderboards.
fable-coder is a chained distill + behavioral fine-tune for Claude-Code-style agentic coding:
Qwen3.6-35B-A3B (Apache-2.0)
└─ Opus-4.7 reasoning distill (lordx64/…-Reasoning-Distilled)
└─ abliteration (huihui-ai/…-abliterated) ← our base
└─ LoRA fine-tune, agentic rounds r3→r4→r6 ← this model (r6)
<think> chains (inherited from the Opus-4.7 prior; intact — verified).This is not a single-teacher distillation from scratch, and it does not aim to exceed its teachers. It is a behavioral graft: the reasoning comes from the Opus-4.7 distill in the base; our LoRA rounds add agentic coding behavior distilled from verified Claude Fable-5 / Opus-4.8 Claude Code sessions. Evaluate and use it accordingly:
| Setting | Value |
|---|---|
| Base | huihui-ai/Huihui-Qwen3.6-35B-A3B-Claude-4.7-Opus-abliterated |
| Method | LoRA (unsloth + TRL SFTTrainer), adapter-continuation across rounds (never restarted from base) |
| LoRA | r=32, α=32, targets = attention (q/k/v/o) + MLP (gate_up_proj, down_proj) |
| Precision | bf16; train_on_responses_only; MoE router aux loss on |
| Seq length | 4096 |
| Final round (r6) | fresh continuation on the r4 adapter, LR 2e-5, 1 epoch (126 steps) |
| r6 data | 1,007 agentic-only windows — verified Claude Code coding sessions (our own generation + community Fable-5 traces + Glint corpus + swarm-salvage). Instruction-pair data from earlier rounds was removed for this round. |
| Data hygiene | rejection-sampled (kept only sessions whose tests/build passed); zero-overlap hash-assert vs prior rounds; secrets/PII scrubbed |
Lineage note (a documented lesson): an intermediate round (r5) that restarted from base with a filtered corpus regressed hard (HumanEval 90.9 → 71.3). The fix — and the method used for r6 — is strict adapter continuation plus an agentic-only final corpus. r6 recovered and improved.
Methodology & honesty. All numbers below are our own harness, q8_0 GGUF, native thinking (temp 0.6 / top-p 0.95), single-sample pass@1, run locally. They are not directly comparable to official leaderboards (different precision, harness, and prompting). AIME uses a 16k-token budget so long reasoning chains don't truncate.
| Benchmark | r6 (this model) | r4 (prior round) | Base (huihui) |
|---|---|---|---|
| HumanEval (pass@1) | 90.2 | 90.9 | 90.2 |
| MBPP (pass@1) | 78.2 | 76.2 | 73.0 |
| GSM8K | 94.7 | 95.0 | — |
| MATH-500 | 88.2 | 89.4 | — |
| AIME 24+25 (16k) | 73.3 | 71.7 | — |
| MMLU-Pro | 79.8 | 77.2 | — |
Read: r6 preserves the base's reasoning (GSM8K/MATH/MMLU-Pro/AIME all healthy) while improving the metric closest to its job — MBPP +5.2 over base, +2.0 over r4 — with no regression on any axis versus r4. Reasoning is preserved; coding — the model's actual job — improves.
🚧 Pending: SWE-bench Lite (agentic harness) is the key remaining test — it measures the actual coding-agent axis these benchmarks can't. Numbers will be added when verified.
Produced locally with llama.cpp from the bf16 master (llama-quantize):
| Quant | Weights | GPU / Mac (with room for context) |
|---|---|---|
| Q8_0 | 38GB | 48GB+ GPU · 64GB Mac — near-lossless |
| Q6_K | 29GB | 40GB+ GPU · 48GB Mac |
| Q5_K_M | 25GB | 32GB GPU |
| Q4_K_M | 22GB | 32GB GPU (or a 24GB card at short context) |
Sizes are the weights only — budget headroom on top for the KV cache + compute buffers. The good news: this model's KV cache is unusually small (only 2 KV heads), so long context is cheap — ~2.7GB at 32k, ~11GB at 128k, ~21GB at the full native 256k. That's why it's comfortable on modest hardware despite being a 35B.
Pre-made GGUF quants (Q4–Q8) → GGUF repo, or ollama run achillessafehavencalls/fable-coder. The full-precision bf16 weights are in this repo — or quantize your own levels (F16, IQ4_XS, etc.) with llama.cpp.
Run it instantly with Ollama:
ollama run achillessafehavencalls/fable-coder
Or serve the GGUFs with llama.cpp / LM Studio / vLLM. Thinking is native — the Qwen template opens <think>
by default; the server returns reasoning in reasoning_content and the answer in content. For
agentic use, run inside a harness that supplies a tool-use system prompt + tool registry (treat it
like Claude Code). Note: tool-name binding is loose at this data scale — downstream tool routers
should normalize invented names (e.g. read_file → Read).
Released under Apache-2.0, consistent with the Qwen3.6-35B-A3B base and the Opus-4.7 distill it builds on (both Apache-2.0). We treat the model weights as an independent artifact, not a derivative work of the training data.
Provenance disclosures (in the spirit of full transparency):
Glint-Research/Fable-5-traces.
Trace contributors are credited under Attribution.Responsible use: this is an uncensored (abliterated-base) coding model released for sovereign/research use. You are responsible for compliance and safety in your deployment. Do not use it to generate malware, conduct unauthorized intrusion, or carry out other unlawful activity.
@misc{fable_coder_35b_2026,
title = {fable-coder-35B-A3B: agentic-coding fine-tune of Qwen3.6-35B-A3B (Claude Fable-5/Opus distill)},
author = {Dappit Labs},
year = {2026},
howpublished = {\url{https://huggingface.co/Achilles1089/fable-coder-35B-A3B}},
}