Downloads · 30 days
242
21% of all-time downloads
RefinedNeuro/RefinedToolCallV5-3b
RefinedToolCallV5-3b is a text generation model from RefinedNeuro. Use it when you need the model to write or continue text. It is set up for hermes. The card lists the license as apache-2.0.
Math-grade reasoning · real function calling · multi-turn agentic · 2.5 GB · runs on your laptop.
Downloads · 30 days
242
21% of all-time downloads
All-time downloads
1.1K
Public
Parameters
3.1B
18.2 GB on disk
Likes
12
Public
Click a slice to open those files.
.gguf12 GB · 66%
From the Hugging Face model README
Math-grade reasoning · real function calling · multi-turn agentic · 2.5 GB · runs on your laptop.
ollama run refinedneuro/refinedtoolcallv5-3b
Most 3B tool-callers nail a single function call and then fall apart the moment the task spans several turns. RefinedToolCall-V5 was built specifically to fix that — and the numbers moved on every axis at once, not just the one we were targeting.
multi_turn) than where we started.| capability | this model |
|---|---|
🔁 Multi-turn agentic (BFCL multi_turn, k=3) | 0.220 avg / 0.298 pass@3 |
| 🛠️ Single-turn function calling (BFCL, held-out) | 0.707 |
| 🩹 Recovery from tool errors (n=250) | 0.896 |
| 🧮 Reasoning (AIME-2024 pass@8) | 0.933 |
Every number is the best across five fine-tuning rounds — multi-turn, single-turn, recovery, and reasoning all peaked together.
We didn't just throw data at it. Five disciplined rounds, each one gated so it could never regress reasoning or recovery:
Ollama
ollama run refinedneuro/refinedtoolcallv5-3b # latest = Q6_K, 2.5 GB
💡 Use Q6_K or higher for tool-calling — lower quants corrupt the call tokens.
Format: ChatML + Hermes tools. Each turn the model emits a <think> plan → one or more
<tool_call> blocks → a final reply. Recommended: temp 0.6, top_p 0.95, repeat_penalty 1.1.
✅ Local/offline agentic tool-use prototypes ✅ Multi-step function-calling assistants ✅ Math & STEM reasoning ✅ Learning how small agentic models are actually built.
⚠️ It's a 3B research preview. Multi-turn is dramatically improved (~3.7×) but not solved — very long, open-ended autonomous loops can still write buggy code or mis-plan. A brilliant, tiny building block; not yet a drop-in autonomous engineer.
Built on WeiboAI/VibeThinker-3B + lambda/hermes-agent-reasoning-traces. Trained with distribution-matched RFT + on-policy expert iteration, every checkpoint gated against reasoning/recovery canaries. Apache-2.0.