Downloads · 30 days
5
28% of all-time downloads
DANIELDX2/bfclv4-submission
bfclv4-submission is a machine learning model from DANIELDX2. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Complete submission repository for Berkeley Function Calling Leaderboard (BFCL) v4.
Downloads · 30 days
5
28% of all-time downloads
All-time downloads
18
Public
Parameters
1.1B
2.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.2 GB · 100%
From the Hugging Face model README
Complete submission repository for Berkeley Function Calling Leaderboard (BFCL) v4.
handler.py - Main entrypoint with handle_request() functionrequirements.txt - Python dependenciesMODEL_CARD.md - Model information and training detailsinference_example.py - Simple usage exampletrain_lora.sh - LoRA fine-tuning script (GPU required)tests/ - Test suite with fixtures.github/workflows/ci.yml - CI pipelinecloud_build_instructions.md - Cloud deployment guidebfcl_submission_instructions.md - Submission checklistMODEL_NAME - HuggingFace model identifier (default: meta-llama/Llama-2-13b-chat)BFCL_MODE - Operation mode: prompt or fc (default: prompt)DEVICE - Compute device: cpu or cuda (default: cpu)HF_TOKEN - HuggingFace token for gated models (optional)Important: Use Python 3.10 or 3.11 (PyTorch doesn't support 3.13 yet)
python3.11 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
Basic usage (uses mock responses for testing):
python inference_example.py
With a real model:
# TinyLlama (2.2GB download)
MODEL_NAME="TinyLlama/TinyLlama-1.1B-Chat-v1.0" DEVICE="cpu" python inference_example.py
# Larger model on GPU
MODEL_NAME="meta-llama/Llama-2-13b-chat" DEVICE="cuda" python inference_example.py
Quick test (uses mock responses):
pytest -q
With real model:
MODEL_NAME="TinyLlama/TinyLlama-1.1B-Chat-v1.0" DEVICE="cpu" pytest -v
Constructs a prompt with available functions and generates output containing FUNCTION_CALL: {"name": "...", "args": {...}}. The handler parses this structured output.
Set BFCL_MODE=fc to use native function calling APIs if supported by the model.
{
"status": "ok",
"response": {
"type": "function_call",
"name": "function_name",
"arguments": {...}
},
"raw_model_output": "..."
}
The handler processes requests that may require multiple function calls. The current implementation returns a single function call per request. Multi-call scenarios are handled through iterative turns.
The test suite covers:
All tests use local fixtures and run deterministically. When no real model is loaded, the system uses intelligent mock responses that extract function information from the available functions list.
LoRA fine-tuning requires GPU:
bash train_lora.sh
Do not run on CPU-only machines.
Apache-2.0