Downloads · 30 days
34
16% of all-time downloads
singulared/Ornith-1.0-35B-MTP-ROCmFP4-GGUF
Ornith-1.0-35B-MTP-ROCmFP4-GGUF is a text generation model from singulared. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as apache-2.0.
⚠️ Requires the ROCmFPX fork of llama.cpp — this will NOT load on mainline llama.cpp. ROCmFP4 is an experimental AMD FP4 quant format (ggml tensor types 100–107) that only exists in that fork. Targets AMD Strix Halo /…
Downloads · 30 days
34
16% of all-time downloads
All-time downloads
208
Public
Repo size
20.6 GB
Likes
0
Public
Click a slice to open those files.
.gguf20.6 GB · 100%
From the Hugging Face model README
⚠️ Requires the ROCmFPX fork of llama.cpp — this will NOT load on mainline llama.cpp. ROCmFP4 is an experimental AMD FP4 quant format (ggml tensor types 100–107) that only exists in that fork. Targets AMD Strix Halo / Radeon 8060S (gfx1151). For standard GGUF that runs anywhere, use the base repo: → singulared/Ornith-1.0-35B-MTP-GGUF (Q8_0 / Q4_K_M, mainline llama.cpp).
Ornith-1.0-35B (DeepReinforce, a qwen35moe
agentic coder) with an embedded MTP (nextn) head grafted in, quantized to ROCmFP4-COHERENT
for fast self-speculative decoding on Strix Halo.
Q4_0_ROCMFP4_COHERENT (4.70 bpw — ROCmFP4 experts + Q6_K token embeddings),
quantized directly from the BF16 original (deepreinforce-ai/Ornith-1.0-35B) — a single clean
quantization pass, no intermediate Q8 (avoids double-quant).blk.40.nextn.* grafted raw from Qwen/Qwen3.6-35B-A3B's
native head (same qwen35moe arch), after quantization — so eh_proj stays at its original Q8_0.
The head's eh_proj precision is the throughput lever, and grafting it raw keeps it high. Ornith is a
barely-shifted fine-tune of Qwen3.6-35B-A3B, so the head transfers cleanly (~88% draft acceptance).Recommended: Vulkan (-dev Vulkan0) — best decode, which dominates a coder's token cost:
n4, -ub 2048) · 72 t/s without MTP · draft acceptance 0.88-ub 2048)-ub 2048)| backend | prefill (pp4096) | decode (MTP n4) |
|---|---|---|
| Vulkan | 1051 | 86.7 |
| ROCm/HIP @ ROCm 7.2 | 876 | 70.5 |
| ROCm/HIP @ ROCm 7.15.0a20260721 nightly (compiled) | 1155 | 68.1 |
therock-dist-linux-gfx1151-7.15.0a20260721
(build 2026-07-21), built against the nightly tarball, not a packaged ROCm. Newer nightlies may differ.llama-server -m ornith-1.0-35b-MTP-ROCmFP4-COHERENT.gguf \
-dev Vulkan0 -fa on -ngl 99 -c 131072 --jinja \
--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.6 --alias ornith
The MTP head is embedded — no separate draft model. --spec-draft-n-max 4 is the MoE sweet spot
(deeper drafts get rejected). Build the fork per its Strix Halo quickstart.
Derivative of two permissively-licensed models; both credited, their licenses apply to their parts:
Quantization: ROCmFP4-COHERENT (ROCmFPX fork). Not affiliated with or endorsed by DeepReinforce, Qwen, or the ROCmFPX author.