Downloads · 30 days
157
100% of all-time downloads
Berhak/Llama-3.1-8B-Function-Calling-Agent
Llama-3.1-8B-Function-Calling-Agent is a text generation model from Berhak. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as llama3.1.
Downloads · 30 days
157
100% of all-time downloads
All-time downloads
157
Public
Repo size
5.3 GB
Likes
0
Public
Click a slice to open those files.
.gguf4.9 GB · 93%
From the Hugging Face model README
<a href="https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct"><img alt="Base" src="https://img.shields.io/badge/Base-Llama--3.1--8B--Instruct-0866ff?style=for-the-badge"></a> <img alt="Method" src="https://img.shields.io/badge/Method-QLoRA%20r%3D32-7c3aed?style=for-the-badge"> <img alt="Quant" src="https://img.shields.io/badge/Quant-Q4__K__M%20%2B%20imatrix-059669?style=for-the-badge"> <br> <img alt="Unseen schemas" src="https://img.shields.io/badge/Unseen%20schemas-92%25-16a34a?style=for-the-badge"> <img alt="VRAM" src="https://img.shields.io/badge/VRAM-5.1%20GB-ea580c?style=for-the-badge"> <img alt="License" src="https://img.shields.io/badge/License-Llama%203.1-475569?style=for-the-badge">
<sub><b>Built with Llama</b></sub>
</div> <hr>A function-calling fine-tune of Llama-3.1-8B-Instruct, trained to read JSON tool schemas at
inference time and emit well-formed <tool_call> blocks — then answer in natural English once the
tool result comes back. It runs on one 8 GB consumer GPU at ~40 tok/s.
| 🎯 Unseen schemas | 🔑 Seen schemas | ✅ JSON validity | 💾 Total VRAM |
|---|---|---|---|
| 24 / 26 | 48 / 56 | 57 / 57 | 5055 MiB |
| 92% | 85% | 100% | of 7774 MiB |
Those first two columns are the whole result. Near-identical performance on schemas the model has never seen means it learned "read the schema, build the call" rather than the tool names it was trained on.
you > Do you remember what my job is?
agent> Your job is that of a software engineer.
· prompt 713 tok 0.8s (860 t/s) · gen 10 tok 0.2s (38 t/s)
you > What is the current population of Japan?
· google_search({'query': 'current population of Japan', 'num_results': 1})
-> 3 sources (wikipedia): en.wikipedia.org
agent> According to Wikipedia, Japan's population peaked at 128.5 million people in 2010
and had declined to approximately 123.4 million by April 2025. By 2050, an estimated
one-third of the population in Japan is expected to be 65 and older due to
sub-replacement fertility rates and high life expectancy. […]
· prompt 1390 tok 1.3s (1044 t/s) · gen 137 tok 3.8s (36 t/s)
you > Do you remember which city I live in? Could you check what the weather will be
like there tomorrow and whether I should take an umbrella?
· get_weather({'location': 'New York', 'unit': 'celsius', 'days_ahead': 1}) -> ok
agent> In New York, there will be rain tomorrow with a minimum temperature of 19.6 degrees
Celsius and a maximum of 25.6 degrees Celsius. It would be wise to bring an umbrella.
· prompt 420 tok 0.3s (1232 t/s) · gen 41 tok 1.1s (37 t/s)
Two things in there are the card's claims, live:
days_ahead was never in the training data. It is an extension to the get_weather schema
added after the fine-tune, and the model reads "tomorrow" → days_ahead: 1 straight out of the
parameter description — the same schema-reading ability the 92% measures.num_results: 1 and got three sources back: the reference agent
overriding a learned habit, see It asks for one search result under Limitations.Two separate artifacts. They are not interchangeable.
| file | what it is | needs |
|---|---|---|
🟢 llama31-8b-Q4_K_M.gguf<br><sub>4.92 GB</sub> | Merged model — base + LoRA, imatrix-calibrated | llama.cpp / llama-server. Self-contained |
🔵 adapter_model.safetensors<br>adapter_config.json<br><sub>336 MB</sub> | The LoRA adapter only | unsloth/Meta-Llama-3.1-8B-Instruct + PEFT |
llama.cpp
./llama-server -m llama31-8b-Q4_K_M.gguf \
-c 8192 -ngl 99 -fa on \
--cache-type-k q8_0 --cache-type-v q8_0 \
--host 127.0.0.1 --port 8080
The q8_0 KV cache halves cache cost to ~68 KB/token, which is what makes 8192 context fit alongside the weights on an 8 GB card.
transformers + PEFT
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "unsloth/Meta-Llama-3.1-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, "Berhak/Llama-3.1-8B-Function-Calling-Agent")
tokenizer = AutoTokenizer.from_pretrained(base)
<hr>
⚠️ This matters more than the weights.
The model was trained on one specific shape and behaves poorly outside it. All three pieces below are load-bearing.
Tool schemas go inside <tools> as a JSON array, in Llama-3.1's native header format:
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 12 Sep 2026
You are a function calling AI model. You are provided with function signatures within <tools></tools> XML tags. You may call one or more functions to assist with the user query. Don't make assumptions about what values to plug into functions.
<tools>
[{"type": "function", "function": {"name": "get_weather", "description": "Get current or forecast weather conditions for a location.", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "The city name, e.g. 'Ankara'."}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "Temperature unit."}}, "required": ["location"]}}}]
</tools>
For each function call return a json object with function name and arguments within <tool_call></tool_call> XML tags.<|eot_id|>
📅 Pass the real date. strftime("%b") follows the OS locale and produces a non-English month
abbreviation under a non-English one, which breaks the template.
<tool_call>
{"name": "get_weather", "arguments": {"location": "Berlin", "unit": "celsius"}}
</tool_call>
🔢 It emits several blocks in one generation when the request needs several tools. Parse every block, not just the first.
Tool results go back on the ipython role, double-encoded: the <tool_response> block is
wrapped in a JSON string, escaped quotes and literal \n included. An artifact of the original
Hermes conversion, but the model learned that exact shape:
<|start_header_id|>ipython<|end_header_id|>
"<tool_response>\n{\"name\": \"get_weather\", \"content\": {\"location\": \"Berlin\", \"temperature\": 18.4}}\n</tool_response>"<|eot_id|>
observation = json.dumps(
"<tool_response>\n" + json.dumps({"name": name, "content": result}) + "\n</tool_response>"
)
<hr>🚨 A plain, single-encoded block puts the model off its training distribution.
Build a GBNF grammar from the tool schemas you pass in this request — not from a fixed file — with a root that permits either a run of calls or plain prose:
root ::= tool-call (ws-nl tool-call)* | prose
tool-call ::= "<tool_call>" ws-nl call ws-nl "</tool_call>"
prose ::= [^<] ( [^<] | "<" )*
ws-nl ::= [ \t\n]*
<div align="center">
| grammar off | grammar on | |
|---|---|---|
| JSON validity | 56/59 · 94% | 57/57 · 100% |
| Overall | 71/82 · 86% | 72/82 · 87% |
| Unparseable blocks | 3 | 0 |
No category regressed, and the gain concentrates where the model is weakest unaided — generations needing more than one call. Three details are worth copying verbatim:
prose bans < only in the first position. Banning it everywhere ([^<]+) made the
character unrepresentable — asked for an inequality the model wrote 3 ≤ 5 instead of 3 < 5,
a different claim, not a formatting quirk.86 hand-written held-out records, screened against the training data for verbatim matches and at a 0.70 similarity threshold. Every tool name claimed as "unseen" was checked against all 2,985 tool names present in training. Scored with the grammar on:
| metric | score | |
|---|---|---|
| 🎯 Unseen tool schemas | 24 / 26 | █████████░ 92% |
| 🔑 Seen tool schemas | 48 / 56 | ████████░░ 85% |
| ✅ JSON validity | 57 / 57 | ██████████ 100% |
| 🎓 Calls built correctly from an unseen schema | 11 / 11 | ██████████ 100% |
| 🚫 Not calling a tool when none is needed | 8 / 8 | ██████████ 100% |
| 🌐 Live-information questions → search | 6 / 6 | ██████████ 100% |
| 🧮 Arithmetic routed to a calculator tool | 5 / 5 | ██████████ 100% |
| 🔗 Two calls in one generation | 2 / 2 | ██████████ 100% |
| 📋 Several tasks in one message | 5 / 7 | ███████░░░ 71% |
| 🪤 Keyword collision without picking the wrong tool | 3 / 5 | ██████░░░░ 60% |
| 👻 Not fabricating a missing required argument | 1 / 4 | ██░░░░░░░░ 25% |
| 🧩 Overall | 72 / 82 | █████████░ 87% |
Four ambiguous records are reported but never scored — both behaviours are defensible there, and inventing a reference answer to complete a metric would only make the metric worse.
A handful of records also score the language of the final answer against Turkish (see below), so those are not an English measurement. Every row above except Overall measures tool selection and argument construction, which is language-independent.
<hr>🚧 It does not call a tool after seeing an observation.
This is structural, not a weak tendency.
Of the 8,144 assistant turns that follow a tool result across 9,146 training examples, zero
are a tool call. The data only ever contains CALL → OBSERVE → ANSWER; a second call appears solely
after a new user turn. A directive telling it to continue scored 0/3.
✅ Plan around it: a request needing several tools must be satisfied by the first generation. The model does this reliably, so parse every block it emits. Do not build a loop that expects it to iterate.
👻 It fabricates arguments it does not have. "What's the weather?" with no city produced
location='Ankara'; "is there a train from Ankara?" produced destination="user's destination" —
a literal placeholder, meaning it knows it does not know and fills the slot anyway. No prompt
variant fixed this. Validate arguments against the conversation before dispatching.
1️⃣ It asks for one search result. The training data calls search with num_results=1 in 34 of 55
cases. Treat that parameter as a floor in your tool, not a ceiling.
🎭 It invents explanations for things that do not exist. Asked about a fabricated syndrome, it answers confidently and fluently. Empty or failed tool results should say so explicitly in the observation, and say what the model must not do.
🎯 Search queries drift off subject. Asked to search two topics, the second query sometimes lands on a neighbouring one — worse with accumulated context.
🧮 Arithmetic needs a tool. Unaided it produced "1 FP16 = 2 INT4, so 16 GB becomes 32 GB" — the ratio is 4 and the operation is division — then reused the wrong figure as a premise on the next turn. Route calculations to a calculator tool.
📏 Factual accuracy of direct answers was never measured, and 📉 no bfloat16 baseline exists — the merged model was deleted by the quantization pipeline before one was taken, so every number here is absolute rather than a measurement of quantization loss.
🌍 On Turkish. The training mix includes a small Turkish subset (~5%), and the evaluation set was originally written against Turkish final turns. That path is not validated and is not recommended — the agent around this model was brought to a working standard in English only. Treat this as an English function-calling model.
<hr>QLoRA on 4× H100, a single run with no retry budget. Source:
NousResearch/hermes-function-calling-v1, converted from ChatML into Llama-3.1's native chat
template before loss masking — Llama-3.1's tokenizer does not recognize <|im_start|> as a special
token, so raw ChatML would shatter into meaningless sub-words.
| 🎛️ LoRA | r=32, α=64, dropout 0.05, 4-bit NF4 + double quant, bf16 compute |
| 🎯 Target modules | q_proj k_proj v_proj o_proj gate_proj up_proj down_proj |
| 📏 / 🔁 | 4096 tokens · 2 epochs, early stopping (patience 3) |
| 📉 LR | 2e-4 cosine, 3% warmup, weight decay 0.01 |
| 📦 Batch | 2 per device × 8 gradient accumulation |
| 🗃️ Data | 9,146 examples → 8,781 train / 365 eval |
🕳️ One warning worth repeating. English final-turn prose was masked out so it would not compete with the intended final-turn behaviour. But Hermes examples that answer without a tool consist of nothing but a final turn — so masking it masked the whole example, and 803 of 851 vanished silently. The resulting skew toward calling a tool is why the model over-triggers, and why the reference agent adds a restraint clause to the system prompt. If you reuse this recipe: count your examples after masking, not before.
Quantization: bf16 merge → GGUF q8_0 → imatrix → Q4_K_M. The intermediate is q8_0 rather than f16
(half the disk, no measurable K-quant penalty; needs --allow-requantize), and the importance
matrix was calibrated on 40 samples drawn from the real training distribution.
RTX 4060 Laptop (8 GB), Vulkan, n_ctx=8192, Q4_K_M, q8_0 KV cache, flash attention:
| 🧠 Weights | 4403 MiB |
| 🗂️ KV cache | 544 MiB · 68 KB/token |
| ⚙️ Compute buffer | 108 MiB |
| 💾 Total | 5055 / 7774 MiB |
| ✍️ Generation | ~40 tok/s |
| 📥 Prompt processing | ~175 tok/s |
🚦 Memory is not the binding constraint — prompt processing is.
At 175 tok/s, every 1000 tokens of context costs about six seconds on every subsequent turn. Trim tool output aggressively and keep the system prefix stable so prefix caching holds (measured: 14 tokens reprocessed instead of 440 on the second request).
Sampling: temperature=0.5 top_k=40 top_p=0.9 repeat_penalty=1.1 repeat_last_n=768.
At temperature=0 the model locked into repeating one sentence across turns; repeat_penalty is
what broke that, not the temperature. Use temperature=0 for reproducible evaluation.
Llama 3.1 Community License. Use is subject to the Llama 3.1 License and the Acceptable Use Policy.
Built with Llama.
Training data: NousResearch/hermes-function-calling-v1,
subject to its own terms.
@misc{llama31-8b-function-calling-agent,
title = {Llama-3.1-8B-Function-Calling-Agent},
author = {Tanyıldızı, Mahmut Berhak},
year = {2026},
url = {https://huggingface.co/Berhak/Llama-3.1-8B-Function-Calling-Agent}
}