Downloads Β· 30 days
201
16% of all-time downloads
hotdogs/qwen27b-agent-R2-preview
qwen27b-agent-R2-preview is a text generation model from hotdogs. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as agpl-3.0.
<h1 align="center"π qwen27b-agent-R2-preview</h1
Downloads Β· 30 days
201
16% of all-time downloads
All-time downloads
1.3K
Public
Parameters
26.9B
298 GB on disk
Likes
0
Public
Click a slice to open those files.
.gguf94.8 GB Β· 63%
From the Hugging Face model README
Preview release β Built on Qwen3.6-27B with multi-LoRA fusion. Features Multi-Token Prediction (MTP) for speculative decoding, tool-calling, and Opus + Fable reasoning. Standard (non-abliterated) version.
| Capability | Description |
|---|---|
| β‘ MTP Speculative Decoding | Draft 2 tokens at a time β up to +85% decode TPS on single GPU |
| π§ Tool Calling | Hermes/Qwen function-calling format via llama.cpp --tools all |
| π§ Reasoning | Opus 4.8 + Fable-style reasoning with step-by-step CoT |
| π Thai + English | Native bilingual support |
| π» Code | Python, shell, system tasks |
# Quick test
./llama-cli -m qwen27b-agent-R2-preview.Q4_K_M.gguf \
-p "Hello" -n 100 --temp 0.6
# Full agent server with tool calling + MTP speculative decoding
./llama-server \
-m qwen27b-agent-R2-preview.Q4_K_M.gguf \
--host 0.0.0.0 \
--port 8081 \
-c 262144 \
-ngl 99 \
--cache-type-k bf16 \
--cache-type-v bf16 \
--flash-attn on \
--tools all \
--cont-batching \
--temp 0.6 \
--top-k 40 \
--top-p 0.9 \
--min-p 0.05 \
--repeat-penalty 1.03 \
--dry-multiplier 0 \
--verbose \
-n -1 \
--parallel 1 \
--jinja \
--dry-sequence-breaker none \
--spec-type draft-mtp \
--spec-draft-n-max 2
| Parameter | Purpose |
|---|---|
--cache-type-k bf16 / --cache-type-v bf16 | BF16 KV cache for quality |
--flash-attn on | Flash attention for speed |
--tools all | Enable tool/function calling |
--spec-type draft-mtp | MTP speculative decoding (draft 2 tokens) |
--spec-draft-n-max 2 | Max 2 draft tokens per step |
--cont-batching | Continuous batching for multi-turn |
--jinja | Use Jinja2 chat template from GGUF |
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"hotdogs/qwen27b-agent-R2-preview",
torch_dtype="auto",
device_map="auto",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("hotdogs/qwen27b-agent-R2-preview")
messages = [{"role": "user", "content": "Hello"}]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, return_tensors="pt")
outputs = model.generate(inputs, max_new_tokens=256, temperature=0.6)
print(tokenizer.decode(outputs[0]))
| File | Size | Quant | Description |
|---|---|---|---|
qwen27b-agent-R2-preview.Q4_K_M.gguf | 16 GB | Q4_K_M | Recommended β balanced quality/speed |
qwen27b-agent-R2-preview.Q6_K.gguf | 21 GB | Q6_K | Higher quality, slightly slower |
qwen27b-agent-R2-preview.f16.gguf | 51 GB | f16 | Full precision |
π― Q4_K_M is recommended for most users β good quality with 16 GB VRAM usage.
For vision support, pair this model with the mmproj from Qwen/Qwen3.6-27B:
# Extract mmproj from Qwen3.6-27B vision model
python3 ./llama.cpp/convert_hf_to_gguf.py \
--mmproj Qwen/Qwen3.6-27B \
--outfile mmproj-qwen3.6-27b.gguf
# Use with llama-server for vision + tool calling
./llama-server \
-m qwen27b-agent-R2-preview.Q4_K_M.gguf \
--mmproj mmproj-qwen3.6-27b.gguf
| Parameter | Value |
|---|---|
| Base | Qwen/Qwen3.6-27B |
| Parameters | ~27B |
| Hidden Size | 5,120 |
| Attention | Linear + Standard hybrid |
| Context | 8,192 tokens (extendable) |
| Precision | BF16 / GGUF quantized |
| Format | ChatML (Jinja2 template) |
| MTP Head | β 1 extra layer (draft 2 tokens) |
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β qwen27b-agent-R2-preview Construction β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β Qwen/Qwen3.6-27B (Base) β
β β β
β βββ Multi-LoRA Fusion (4 LoRAs): β
β β ββββββββββββββββββββββββββββββββββββ β
β β β LoRa β Source β β
β β ββββββββββββββββΌββββββββββββββ€ β
β β β Opus SFT β SFT on β β
β β β β Opus 4.8 β β
β β ββββββββββββββββΌββββββββββββββ€ β
β β β CxCMU Agent β AgentWorld β β
β β β β trajectories β β
β β ββββββββββββββββΌββββββββββββββ€ β
β β β General SFT β Reasoning β β
β β β β + Hermes FC β β
β β ββββββββββββββββΌββββββββββββββ€ β
β β β Tachibana β Coding agent β β
β β β Agent β dataset β β
β β ββββββββββββββββ΄ββββββββββββββ β
β β β
β βββ + MTP Head (15 tensors) β
β βββ From huihui-ai/Huihui-Qwen3.6-27B-abliterated β
β β
β Result: 866 tensors, MTP=1, 16.8 GB (Q4_K_M) β
β β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Each LoRA was trained independently via SFT on specialized datasets. Then all 4 were fused into the base model at their respective scales. The low scales (0.15-0.4) ensure no single LoRA overpowers the model β creating a balanced agent.
The Multi-Token Prediction head (15 tensors) was injected from huihui-ai/Huihui-Qwen3.6-27B-abliterated to enable speculative decoding:
# MTP head adds blk.64.* tensors for draft-2-token prediction
--spec-type draft-mtp # Enables in llama.cpp
--spec-draft-n-max 2 # Max draft tokens
Multi-Token Prediction enables speculative decoding:
Standard: [tokenβ] β [tokenβ] β [tokenβ] β ... (~36 TPS)
MTP: [tokenβ tokenβ] β [tokenβ tokenβ] β ... (~66 TPS)
--spec-type draft-mtp in llama.cppIf you find this model useful, please consider supporting my work!
ΰΈ«ΰΈ²ΰΈΰΈΰΈΈΰΈΰΈΰΈ΄ΰΈΰΈ§ΰΉΰΈ²ΰΉΰΈ‘ΰΉΰΈΰΈ₯ΰΈΰΈ΅ΰΉΰΈ‘ΰΈ΅ΰΈΰΈ£ΰΈ°ΰΉΰΈ’ΰΈΰΈΰΉ ΰΈΰΈ£ΰΈΈΰΈΰΈ²ΰΈͺΰΈΰΈ±ΰΈΰΈͺΰΈΰΈΈΰΈΰΈΰΈ₯ΰΈΰΈ²ΰΈΰΈΰΈΰΈΰΈΰΈ±ΰΈΰΈΰΉΰΈ§ΰΈ’ΰΈΰΈ°ΰΈΰΈ°! π
bc1qf27cyk3vmugcdyv9xdtuv5jwz37863crpj5c9v
Thank you for your support! πβ¨
ΰΈΰΈΰΈΰΈΰΈΈΰΈΰΈ‘ΰΈ²ΰΈΰΉ ΰΈͺΰΈ³ΰΈ«ΰΈ£ΰΈ±ΰΈΰΈΰΈ²ΰΈ£ΰΈͺΰΈΰΈ±ΰΈΰΈͺΰΈΰΈΈΰΈΰΈΰΉΰΈ²! ππ€
Built with β€οΈ by UKA β 18-year-old coder & cybersecurity expert