Downloads · 30 days
14
25% of all-time downloads
79Labs/astraforge-70b-TCR
astraforge-70b-TCR is a text generation model from 79Labs. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as llama3.3.
Developed by 79Labs · Version 1.0.0 · Changelog
Downloads · 30 days
14
25% of all-time downloads
All-time downloads
56
Public
Repo size
846 MB
Likes
0
Public
Click a slice to open those files.
.safetensors829 MB · 98%
From the Hugging Face model README
Developed by 79Labs · Version 1.0.0 · Changelog
astraforge-70b-TCR (Tool-Calling / Retrieval) is a LoRA adapter for
meta-llama/Llama-3.3-70B-Instruct specialised for reliable agentic tool use: retrieving the
right tool from a large catalog (RAG), eliciting missing parameters, confirming before acting,
emitting schema-valid calls, and staying grounded — while preserving the base model's reasoning. It
is not a general capability upgrade; it is a focused, measurable improvement on the agentic behaviours
that make a tool-using assistant trustworthy in production.
What the evidence supports (and what it doesn't). On an in-house agentic benchmark this model matches the base on reasoning (GSM8K) while substantially improving tool-calling correctness and learning a confirm-before-call discipline the base lacks. The full results table, run log, and machine-readable scores are included in
benchmarks/so every number here is verifiable. Sample size is stated (N=100); treat small-N differences as indicative, not leaderboard-grade.
Trained across 35 business domains (sales/CRM, finance, support, logistics, healthcare, HR, IT, …) on these agentic behaviours:
Use it for building tool-using / function-calling agents where reliability of the agentic protocol (right tool, ask-then-act, confirm, don't hallucinate tools) matters — especially private / on-prem deployments where a hosted frontier API isn't an option.
Do not expect frontier general intelligence. This is a 70B open model specialised on a narrow skill set. For open-ended coding or research, use a larger / code-specialised model.
Comparable open models under identical settings: 4K context (max_seq_len=4096, also AstraForge's
trained/served window), greedy decoding, 2048-token budget so reasoning models finish, answer extracted
after any </think> and from \boxed{} where present, N = 100. Each model is prompted in its own
native tool format (its tokenizer chat template) so none is penalised for a foreign format.
| Model | Reasoning (GSM8K) | Tool-correct | Confirmed-first |
|---|---|---|---|
| astraforge-70b-TCR (this model) | 0.93 | 0.81 | 0.94 |
| Llama-3.3-70B-Instruct (base) | 0.93 | 0.54 | 0.00 |
| Qwen3-32B | 0.80 | 0.79 | 0.02 |
| gpt-oss-120b | 0.87 | 0.43 | 0.05 |
| gpt-oss-20b | 0.87 | 0.46 | 0.06 |
| Gemma-4-31B-it | — | — | — (excluded: generation hang under Unsloth; did not complete) |
Δ vs base: reasoning +0.00, tool-correct +0.27, confirmed-first +0.94.
Evidence: benchmarks/nway_results.json,
benchmarks/nway_run.log,
methodology in benchmarks/BENCHMARK_METHODOLOGY.md.
Reading it honestly:
BFCL v4 simple_python, Prompt mode, N=400: 51.00% (204/400).
One correction, because the earlier internal summary of this run was misleading and we would rather
say so than quietly restate it: an aggregate of "0.43% overall" was computed at the time. That
number averages eleven test categories that were never generated as zero. Only simple_python
was ever run, and on it the model scored 51.00%. The aggregate was an artifact of the harness, not a
measurement of the model.
Function-calling (FC) mode genuinely does score 0.00% — a tool-template defect on our side, still open. The remaining BFCL categories, τ-bench and API-Bank are still unrun.
For comparison on the identical 400 cases, our 8B sibling
(79Labs/astraforge-8b-TCR) scores 37.50%, with
198 of its 250 misses being parse failures rather than wrong tool choices.
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "meta-llama/Llama-3.3-70B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, device_map="auto", load_in_4bit=True)
model = PeftModel.from_pretrained(model, "79Labs/astraforge-70b-TCR")
tok = AutoTokenizer.from_pretrained("79Labs/astraforge-70b-TCR")
tools = [{
"type": "function",
"function": {
"name": "book_flight",
"description": "Book a flight for a traveler.",
"parameters": {"type": "object",
"properties": {"traveler_name": {"type": "string"}, "origin": {"type": "string"},
"destination": {"type": "string"}, "depart_date": {"type": "string"}},
"required": ["traveler_name", "origin", "destination", "depart_date"]}}}]
msgs = [{"role": "user", "content": "Book Ada a flight from SFO to JFK on 2026-08-01."}]
ids = tok.apply_chat_template(msgs, tools=tools, add_generation_prompt=True, return_tensors="pt").to(model.device)
print(tok.decode(model.generate(input_ids=ids, max_new_tokens=256)[0][ids.shape[1]:], skip_special_tokens=True))
With Unsloth:
FastLanguageModel.from_pretrained("79Labs/astraforge-70b-TCR", load_in_4bit=True).
max_seq_len=4096, LR 1e-5, effective batch 16, Unsloth gradient checkpointing.load_best_model_at_end) — stopped at
convergence, best checkpoint at val loss ≈ 0.1198.Governed by the Llama 3.3 Community License (inherited from the base model).
@misc{astraforge70b_tcr_2026,
title = {astraforge-70b-TCR: Tool-Calling and Retrieval Agent (LoRA on Llama-3.3-70B)},
author = {79Labs},
year = {2026},
note = {Continued SFT on 1M agentic examples; eval-gated. Benchmark harness + raw evidence included.},
url = {https://huggingface.co/79Labs/astraforge-70b-TCR}
}