Downloads · 30 days
220
6% of all-time downloads
TuwaiqAcademy/AISA-AR-FunctionCall-Think
AISA-AR-FunctionCall-Think is a text generation model from TuwaiqAcademy. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as gemma.
This model is the organizer-provided baseline for Track B — Reasoning-Augmented Function Calling. It defines the reference score that participating systems are expected to beat. It is released for reproducibility and…
Downloads · 30 days
220
6% of all-time downloads
All-time downloads
3.8K
Public
Parameters
268M
574 MB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors536 MB · 93%
How the weights are stored.
BF16268M · 100%
From the Hugging Face model README
This model is the organizer-provided baseline for Track B — Reasoning-Augmented Function Calling. It defines the reference score that participating systems are expected to beat. It is released for reproducibility and as a starting point — it is not a competition entry.
A compact (270M-parameter) Arabic function-calling model that, given an Arabic user query (in any of 5 dialects) and a set of candidate tools, writes a short Arabic <think> reasoning trace and then emits a structured tool call. Fine-tuned (LoRA) from google/gemma-3-270m on the AISA-ArabicFC reasoning data.
For the non-reasoning Track A baseline, see the sibling model AISA-AR-FunctionCall-FT.
| Role | Official baseline — Track B (Reasoning-Augmented) |
| Base model | google/gemma-3-270m (270M params) |
| Adaptation | LoRA fine-tune (merged), then full causal-LM inference |
| Languages | Arabic — MSA, Gulf, Egyptian, Levantine, Maghrebi |
| Behaviour | <think> Arabic reasoning → structured function call |
| Training data | TuwaiqAcademy/AISA-ArabicFC |
| License | Gemma (see License below) |
Given an Arabic user query and a set of candidate tool definitions, a system must:
<think> … </think>) before the call.| Track | Description |
|---|---|
| A — Core | Decide / Select / Extract |
| B — Reasoning-Augmented ← this model | Track A + an Arabic <think> reasoning trace |
| C — Cross-Dialect Robustness | Diagnostic: dialect-stratified evaluation of A/B submissions |
This model uses Gemma 3 chat turns with a custom function-calling schema (it does not emit plain JSON). The exact prompt is the text field in the dataset; the structure is:
<bos><start_of_turn>developer
<system instruction in Arabic>
<start_function_declaration>declaration:NAME{description:<escape>…<escape>,parameters:{…}}<end_function_declaration>
…one declaration per candidate tool…<end_of_turn>
<start_of_turn>developer
التاريخ والوقت الحالي …: 2024-04-12T23:05:24
اليوم هو الجمعة
أنت نموذج يمكنه استدعاء الوظائف التالية<end_of_turn>
<start_of_turn>user
أريد مقارنة أسعار تلفاز سامسونج في الأردن<end_of_turn>
<start_of_turn>model
The model then generates:
<think>
يبدو أن نية المستخدم هي الحصول على مقارنة لأسعار تلفاز سامسونج في الأردن. أداة "compare_prices" هي الأنسب …
</think>
<start_function_call>call:compare_prices{country:<escape>Jordan<escape>,product_name:<escape>Samsung TV<escape>}<end_function_call>
For a query that needs no tool, the model omits the <start_function_call> block (→ requires_function = false).
import re, torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "TuwaiqAcademy/AISA-AR-FunctionCall-Think"
tok = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID, torch_dtype=torch.float32, device_map="auto"
).eval()
def parse_model_output(text: str) -> dict:
"""Turn raw generation into the shared-task submission schema."""
out = {"requires_function": False, "function_name": "none", "arguments": {}, "think": ""}
if (m := re.search(r"<think>\s*(.*?)\s*</think>", text, re.DOTALL)):
out["think"] = m.group(1).strip()
if (m := re.search(r"<start_function_call>\s*call:(\w+)\{(.*?)\}\s*<end_function_call>", text, re.DOTALL)):
out["requires_function"] = True
out["function_name"] = m.group(1)
for key, str_val, num_val in re.findall(r"(\w+):(?:<escape>(.*?)<escape>|([^,}]+))", m.group(2)):
val = str_val if str_val else num_val
try:
val = float(val) if "." in str(val) else int(val)
except (ValueError, TypeError):
pass
out["arguments"][key] = val
return out
# Easiest path: take the ready-made prompt from the dataset's `text` field and
# cut it at the model turn (everything after is what the model should produce).
from datasets import load_dataset
row = load_dataset("TuwaiqAcademy/AISA-ArabicFC", split="validation")[0]
prompt = row["text"].split("<start_of_turn>model\n")[0] + "<start_of_turn>model\n"
inputs = tok(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
with torch.no_grad():
gen = model.generate(**inputs, max_new_tokens=250, do_sample=False) # greedy
raw = tok.decode(gen[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False)
print(parse_model_output(raw))
# → {'requires_function': True, 'function_name': 'compare_prices',
# 'arguments': {'country': 'Jordan', 'product_name': 'Samsung TV'},
# 'think': 'يبدو أن نية المستخدم …'}
The parsed dict maps directly onto a leaderboard submission line: {"id", "tool_called", "arguments", "think"} (use function_name → tool_called).
Scored on the AISA-ArabicFC held-out test set (1,000 positive + negative examples) using the official v2 metrics:
none)<think> trace0.40·FnAcc + 0.60·ArgEM0.30·FnAcc + 0.50·ArgEM + 0.20·ThinkRate| System | FnAcc | ArgEM | Overall (A) | Overall (B) |
|---|---|---|---|---|
| AISA-AR-FunctionCall-Think (270M) ← this | 0.982 | 0.541 | 0.717 | 0.739 |
| GPT-4o — zero-shot | 0.927 | 0.070 | 0.413 | 0.313 |
| GPT-4o — 3-shot | 0.854 | 0.122 | 0.415 | 0.317 |
| Random baseline | 0.047 | 0.033 | 0.039 | 0.031 |
Key takeaways
google/gemma-3-270m<think> tracesIntended use
Out of scope / limitations
@inproceedings{najar2026aisaarabicfc,
title = {AISA-ArabicFC: Arabic Function Calling for Agentic AI Systems},
author = {Najar, Omar},
booktitle = {Proceedings of the Fourth Arabic Natural Language Processing Conference (ArabicNLP 2026)},
year = {2026}
}
This model is a derivative of Gemma 3 and is distributed under the Gemma Terms of Use. By using it you agree to those terms and to the Gemma Prohibited Use Policy. The AISA-ArabicFC dataset is released separately under Apache-2.0.
Shared-task organizers — [email protected] · Tuwaiq Academy