Downloads · 30 days
9
13% of all-time downloads
whichcy/llama3.2-1B-tool
llama3.2-1B-tool is a text generation model from whichcy. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.2.
This is a tool-calling SFT model based on meta-llama/Llama-3.2-1B-Instruct. It is intended for research and experimentation with single-turn function calling, tool selection, and structured argument generation.
Downloads · 30 days
9
13% of all-time downloads
All-time downloads
69
Public
Parameters
1.2B
7.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.9 GB · 100%
From the Hugging Face model README
This is a tool-calling SFT model based on meta-llama/Llama-3.2-1B-Instruct. It is intended for research and experimentation with single-turn function calling, tool selection, and structured argument generation.
The model was fine-tuned on a processed version of Team-ACE/ToolACE. The training data keeps single-turn tool-use examples and follows the Llama chat format.
meta-llama/Llama-3.2-1B-InstructLlamaForCausalLMThe model uses Team-ACE/ToolACE after lightweight preprocessing:
The processed dataset contains 8,626 examples before the final train/evaluation split.
This repository uses a Llama 3.2 model, so the RoPE configuration is important for correct generation. Different Transformers and inference-server versions may expect different config field names.
For newer Transformers exports, the equivalent configuration is usually represented as:
{
"rope_parameters": {
"factor": 32.0,
"high_freq_factor": 4.0,
"low_freq_factor": 1.0,
"original_max_position_embeddings": 8192,
"rope_theta": 500000.0,
"rope_type": "llama3"
},
"dtype": "float32"
}
For older Transformers, SGLang, vLLM, or BFCL-style evaluation stacks, the same values may need to be represented as:
{
"rope_scaling": {
"factor": 32.0,
"high_freq_factor": 4.0,
"low_freq_factor": 1.0,
"original_max_position_embeddings": 8192,
"rope_type": "llama3"
},
"rope_theta": 500000.0,
"torch_dtype": "float32"
}
The uploaded model uses the second form for broader compatibility with BFCL/SGLang-style inference. This changes config metadata only; it does not change the model weights.
The model was evaluated on BFCL v4 and compared with the base Llama-3.2-1B-Instruct-FC model.
| Task | Llama-3.2-1B-Instruct-FC | LLaMa-3.2-1B-Tool | Improvement |
|---|---|---|---|
parallel_multiple | 15.00 | 72.00 | +57.00 |
simple_python | 73.75 | 88.25 | +14.50 |
parallel | 42.50 | 79.00 | +36.50 |
simple_java | 16.00 | 62.00 | +46.00 |
multiple | 51.00 | 87.50 | +36.50 |
simple_javascript | 40.00 | 66.00 | +26.00 |
irrelevance | 36.25 | 78.33 | +42.08 |
live_irrelevance | 67.53 | 66.97 | -0.56 |
live_parallel_multiple | 0.00 | 45.83 | +45.83 |
live_multiple | 7.12 | 53.47 | +46.35 |
live_parallel | 0.00 | 56.25 | +56.25 |
live_simple | 30.62 | 60.08 | +29.46 |
live_relevance | 43.75 | 100.00 | +56.25 |
| Average | 32.66 | 70.90 | +38.24 |
Overall, the fine-tuned model improves the BFCL v4 average score from 32.66 to 70.90, with the largest gains on parallel, multi-tool, and live relevance tasks. The only category without improvement is live_irrelevance, where performance is nearly unchanged.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "whichcy/llama3.2-1B-tool"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
functions_json = """[
{
"name": "get_weather",
"description": "Get the weather for a location.",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "The city and region."},
"date": {"type": "string", "description": "The date for the forecast."}
},
"required": ["location", "date"]
}
}
]"""
system_prompt = f"""You are an expert in composing functions. You are given a question and a set of possible functions. Based on the question, you will need to make one or more function/tool calls to achieve the purpose.
If none of the functions can be used, point it out. If the given question lacks the parameters required by the function, also point it out.
You should only return the function calls in your response.
If you decide to invoke any of the function(s), you MUST put it in the format of [func_name1(params_name1=params_value1, params_name2=params_value2...), func_name2(params)]
You SHOULD NOT include any other text in the response.
At each turn, you should try your best to complete the tasks requested by the user within the current turn. Continue to output functions to call until you have fulfilled the user's request to the best of your ability. Once you have no more functions to call, the system will consider the current turn complete and proceed to the next turn or task.
Here is a list of functions in JSON format that you can invoke.
{functions_json}
"""
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": "Find the weather in Shanghai tomorrow."},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=256,
do_sample=False,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=False))
When using this model for tool calling, include the tool schemas in either the system prompt or the user prompt, depending on your inference framework. The BFCL-style prompt places the available tools in the system prompt.
Please cite the Llama 3.2 base model and the ToolACE dataset if you use this model in research.