Downloads · 30 days
534
100% of all-time downloads
prithivMLmods/Smaug-Mini-GGUF
Smaug-Mini-GGUF is a image-text-to-text model from prithivMLmods. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Smaug-Mini is Abacus.AI's agentic finetune of Qwen3.8-27B, trained via on-policy reinforcement learning (GRPO, LoRA-merged into the language trunk only, vision tower left bitwise-identical to the base) over multi-turn…
Downloads · 30 days
534
100% of all-time downloads
All-time downloads
534
Public
Repo size
90.5 GB
Likes
1
Public
Click a slice to open those files.
.gguf90.5 GB · 100%
From the Hugging Face model README
Smaug-Mini is Abacus.AI's agentic finetune of Qwen3.8-27B, trained via on-policy reinforcement learning (GRPO, LoRA-merged into the language trunk only, vision tower left bitwise-identical to the base) over multi-turn, tool-using automation episodes with verified, outcome-based rewards, targeting more reliable end-to-end tool use and automation performance while holding general capabilities at parity with the base model. It retains Qwen3.8-27B's architecture, layout, 262,144-token context, and
xhigh/medium/lowreasoning-effort interface exactly, functioning as a drop-in replacement, while delivering notable agentic gains — +4.5 on AutomationBench (41.8 vs. 37.3), +17.1 on JobBench (50.5 vs. 33.4), +13.5 on NL2Repo-Bench (55.8 vs. 42.3), and +2.0 overall on LiveBench (76.9 vs. 75.3) — alongside modest improvements on reasoning benchmarks like HLE and IFBench, with GPQA-diamond and MMMU-Pro held essentially at parity. Its key behavioral shift is redistributing deliberation rather than adding it: episodes finish about three steps sooner at roughly unchanged total reasoning volume, and episodes that exhaust their step budget without completing drop from 3.4% to 1.0%. It's served via vLLM with a Qwen3 reasoning parser and tool-call parser (recommended sampling: temperature 1.0, top_p 0.95, reasoning effortxhigh), with the inherited MTP head left untrained against the updated trunk — so speculative decoding via MTP should stay disabled — and is released under Apache 2.0, inherited from Qwen3.8-27B.
| File Name | Quant Type | File Size | File Link | Description |
|---|---|---|---|---|
| Smaug-Mini.BF16.gguf | BF16 | 53.8 GB | Link | Full BF16 weights. Highest quality, largest file size. |
| Smaug-Mini.Q4_K_M.gguf | Q4_K_M | 16.5 GB | Link | Good quality, default size for most use cases, recommended. |
| Smaug-Mini.Q5_K_M.gguf | Q5_K_M | 19.2 GB | Link | High quality, recommended. |
| Smaug-Mini.mmproj-bf16.gguf | mmproj-bf16 | 931 MB | Link | Multimodal projection file in BF16 format. Used for vision/language models. |
LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp