Downloads · 30 days
118
31% of all-time downloads
MergeAILab/Merge-27B-MTP-Reasoning-v1
Merge-27B-MTP-Reasoning-v1 is a text generation model from MergeAILab. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
Downloads · 30 days
118
31% of all-time downloads
All-time downloads
379
Public
Parameters
27.4B
129 GB on disk
Likes
0
Public
Click a slice to open those files.
.gguf73.3 GB · 57%
From the Hugging Face model README
Merge AI Lab · dense reference model
A LoRA supervised fine-tune of Unsloth's Qwen3.6-27B-MTP, trained locally on a single AMD Radeon 8060S (ROCm) at Merge. Multi-token-prediction (MTP) heads and the vision tower are preserved from the base; only the language tower is adapted.
This is one of two models we are publishing from the same internal programme. It is the dense reference — the quality bar that our Mixture-of-Experts sibling, Merge-35B-A3B-Reasoning-v5-b1, was built to chase. Both were developed against the same capability targets and scored on the same internal benchmark.
We are releasing these because they are honest, reproducible results from a small lab running on one consumer GPU, and because the recipe — what moved the score and what did not — is more useful shared than kept.
Both Merge models were developed and evaluated against one capability set, drawn from the work our internal agents actually do:
| Capability | What that means here |
|---|---|
| Reasoning | Long-form deliberate reasoning with a real thinking budget, then a committed answer. |
| Agentic / tool calling | Multi-step tool use: calling, reading results back, and not re-calling a tool it has already answered from. |
| Coding | Modern web production — ReactJS, Next.js, TypeScript, PostgreSQL (schema, queries, typed server routes). |
| pt-PT creative copy | European Portuguese marketing and interface copy — pt-PT, not pt-BR — with structure discipline, not just translation. |
The end purpose is internal: both models drive MCP servers and tool-using agents across our copy, design and implementation pipeline. They are meant to be the local model behind an agent that plans a page, writes the copy in pt-PT, and implements the component — not a general-purpose chat assistant.
The SFT corpus itself is reasoning-trace distillation (below). The capability list above is what the model was selected and measured against, not a description of the training mix.
Out of scope: anything safety-critical. This fine-tune adds no safety training of its own — the base model's alignment is all that is present.
Scored on our private internal benchmark (described below). Both rows are the same quantization, the same harness, the same grader.
| Model | Quant | Technical /100 |
|---|---|---|
unsloth/Qwen3.6-27B-MTP (base) | Q4_K_M | 89 |
Merge-27B-MTP-Reasoning-v1 | Q4_K_M | 94 |
+5 points over the base, and rank 1 of +50 local models on our rolling table — ahead of its own MoE sibling and every other local build we have tested.
| Secondary metric | Score |
|---|---|
| Vision bonus | 11 / 12 |
| Tool-grounding bonus | 16 / 19 |
A second independent run scored 93 — a 1-point run-to-run spread.
We are not publishing the benchmark, but here is what it is and is not.
It is an 8-task suite built from real briefs in our own production workflow — not synthetic puzzles and not sampled from any public set. Each task is graded against a hidden rubric with withheld answer keys, and the tasks cover exactly the capability list above: design decomposition, planning a content structure from a thin brief, writing and adapting frontend components, pt-PT localization quality, surgical CSS repair from visual intent, client-side React/Next interaction, and a typed Next.js server route backed by PostgreSQL.
Scoring:
Deliberately, the technical score judges output quality, not protocol compliance: a model is not punished on its main score for clumsy tool orchestration, because a bad plan executed cleanly is still a bad plan.
Honest limits, stated plainly:
temp 0.1, the 35B MoE at temp 0.6. That reflects each model at its best,
and is not a controlled A/B.| Field | Value |
|---|---|
| Method | LoRA SFT (PEFT 0.18.1, TRL) |
| Base | Unsloth Qwen3.6-27B-MTP (BF16, local HF conversion of the MTP build) |
| Trainable params | 116,727,808 (0.42% of total) |
| Rank / alpha / dropout | 16 / 16 / 0 |
| Target modules | q_proj k_proj v_proj o_proj, in_proj_qkv in_proj_z in_proj_a in_proj_b out_proj, gate_proj up_proj down_proj |
| Frozen | MTP heads, router, vision tower |
| LR / scheduler / warmup | 2e-4 / cosine / 5% |
| Epochs / steps | 3 / 384 |
| Batch | 1 × grad-accum 8 (effective ~8) |
| Max length | 1024 tokens, no packing |
| Precision / optimizer | BF16 / adamw_torch |
| Hardware | Radeon 8060S, 137.4 GB GTT (53.8 GB model resident); vision tower on CPU |
| Framework | PyTorch 2.12.0a0+rocm7.12 |
| Duration | ~10h37m |
Packing is disabled deliberately so reasoning traces are never split across samples.
1,078 reasoning conversations combining two public TeichAI datasets — Claude Opus 4.5
(250 examples) and Opus 4.6 (887 examples) — deduplicated and filtered to an 8k-token
budget, split 90/10 train/eval. Plain messages turns; no tools, and no private,
client or company data.
Provenance disclosure: these are distilled traces generated by Claude Opus models. Anyone redistributing or building on this model should check both the TeichAI dataset licenses and Anthropic's terms covering the use of model outputs.
| step 5 | step 380 | |
|---|---|---|
| Loss | 1.184 | 0.474 |
| Token accuracy | 0.717 | 0.851 |
| Entropy | 0.652 | 0.489 |
Smooth convergence, no instability; final grad-norm 0.38.
| File | Size | Notes |
|---|---|---|
model-0000{1..6}-of-00006.safetensors | ~52 GB | BF16 merged weights (HF format), MTP patch included |
Merge-27B-MTP-Reasoning-v1-BF16-0000{1,2}-of-00002.gguf | 55 GB | BF16 GGUF, for re-quantizing yourself. Split in two because the single file exceeds the Hub's 50 GB per-file limit — llama.cpp loads a split GGUF by pointing at the first shard, or use llama-gguf-split --merge to rejoin them. |
Merge-27B-MTP-Reasoning-v1-mtp-Q4_K_M.gguf | 16 GB | the benchmarked build, MTP preserved |
mmproj-F32.gguf | 1.8 GB | vision projector (required for image input) |
lora/ | 467 MB | LoRA adapter, for re-merging onto the base |
Q4_K_M is the only quantization we ship, because it is the only one we benchmarked. If
you want another, convert from the BF16 GGUF — we would rather publish one measured quant
than five unmeasured ones.
Embedded GGUF metadata: general.name = Merged Merge 27b Mtp Reasoning,
general.basename = merged-merge, general.finetune = mtp-reasoning.
These are the settings the 94/100 was produced at. This model is stable at low temperature.
| Parameter | Value |
|---|---|
temperature | 0.1 |
top_k | 40 |
top_p | 0.9 |
min_p | 0.05 |
repeat_penalty | 1.05 |
presence_penalty | off |
The low temperature is deliberate: this model is used for precise technical and structured output, and it does not degrade into repetition there.
If you prefer more varied prose, the Qwen3 reasoning-mode defaults (temp 0.6,
top_p 0.95, top_k 20) also work well, but the reported score was not measured there.
Note: the GGUF embeds temp 1.0, top_p 0.95, top_k 20. Set the values above explicitly
— some runtimes pick up the embedded defaults.
llama.cpp:
llama-server \
-m Merge-27B-MTP-Reasoning-v1-mtp-Q4_K_M.gguf \
--mmproj mmproj-F32.gguf \
--jinja --reasoning-format auto --reasoning-budget 6000 \
--temp 0.1 --top-k 40 --top-p 0.9 --min-p 0.05 --repeat-penalty 1.05 \
-c 124096 -ngl 999 --flash-attn on
The model is trained to emit long reasoning traces; a low --reasoning-budget will
truncate them mid-thought.
Released under Apache 2.0, matching the Qwen3.6-27B base as published by Unsloth
(unsloth/Qwen3.6-27B-MTP-GGUF, unsloth/Qwen3.6-27B — both apache-2.0, derived from
Qwen/Qwen3.6-27B).
Two further conditions apply to anyone building on this model, and they are not ours to grant: the TeichAI dataset licenses, and Anthropic's terms on the use of Claude outputs — the training traces are Claude-distilled. Check both before redistributing.
@misc{merge27b_mtp_reasoning_v1,
title = {Merge-27B-MTP-Reasoning-v1},
author = {Mergeinto.digital},
year = {2026},
note = {LoRA SFT of Unsloth Qwen3.6-27B-MTP on distilled reasoning traces}
}