Downloads · 30 days
24
67% of all-time downloads
79Labs/astraforge-8b-TCR
astraforge-8b-TCR is a text generation model from 79Labs. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as llama3.1.
Developed by 79Labs · Version 2.0.0 (2026-10-05)
Downloads · 30 days
24
67% of all-time downloads
All-time downloads
36
Public
Repo size
353 MB
Likes
0
Public
Click a slice to open those files.
.safetensors168 MB · 90%
From the Hugging Face model README
Developed by 79Labs · Version 2.0.0 (2026-10-05)
This repository was replaced in place. v1.0.0, the earlier adapter with its own model card and benchmarks, stays available under the git tag
v1.0.0:PeftModel.from_pretrained(base, "79Labs/astraforge-8b-TCR", revision="v1.0.0"). v2 uses a different tool-call format (see Prompt format), so update your parser when you upgrade.
astraforge-8b-TCR is a LoRA adapter for meta-llama/Llama-3.1-8B-Instruct. It does three things:
choosing and calling tools from a large catalogue, including many calls in one turn; deciding
when to retrieve (rag_search) and answering from what came back; and reasoning step by
step with a calculator tool.
Every model was prompted in its own native tool format. The Claude and DeepSeek rows were run
through the same harness, parser and splits. Raw outputs are in benchmarks/.
Each cell is first-call accuracy / multi_recall. First-call accuracy means the first tool named is the right one. multi_recall is the share of the reference calls the model actually made. Turns in this catalogue average about 6 calls, so read multi_recall alongside first-call accuracy.
| model | seen_real | heldout_real | heldout_synth | alien_synth | args exact (heldout_real) |
|---|---|---|---|---|---|
| astraforge-8b-TCR v2 | 1.000 / 0.901 | 1.000 / 0.979 | 1.000 / 0.999 | 0.983 / 0.979 | 0.983 |
| Claude (frontier API model) | 0.950 / 0.965 | 0.850 / 0.847 | 0.983 / 0.981 | 0.950 / 0.961 | 0.550 |
| model | row pass | retrieval call decision |
|---|---|---|
| astraforge-8b-TCR v2 | 0.972 | 1.000 |
No external model was run on this set.
| model | no tool | with calculator tool |
|---|---|---|
| astraforge-8b-TCR v2 | 0.706 | 0.900 |
| Claude | 0.997 | 1.000 |
This is the adapter's weak axis. Use the calculator path, and do not rely on unaided multi-step arithmetic. The 9B sibling family (Qwen3.5-based) reaches 0.950 / 1.000 on the same set.
v1's confirmed_first (asks before acting, 0.89), elicitation and GSM8K (0.75) figures came from a
different harness and were not re-run on v2, so this card claims nothing about them. If those
behaviours matter to you, test them, or pin v1.0.0.
v2 is trained on the template in chat_template.jinja, and you must pass
it. Llama-3.1's built-in template rejects assistant turns that carry more than one tool call, and
this model makes several per turn. Tool calls come out as XML:
<atem:function_calls>
<atem:invoke name="get_quote">
<atem:parameter name="symbol">NSE:RELIANCE</atem:parameter>
</atem:invoke>
<atem:invoke name="get_news">
<atem:parameter name="query">Reliance Industries</atem:parameter>
</atem:invoke>
</atem:function_calls>
Parse the invoke blocks and coerce each parameter to its schema type.
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
import torch
base = "meta-llama/Llama-3.1-8B-Instruct"
tok = AutoTokenizer.from_pretrained("79Labs/astraforge-8b-TCR") # loads chat_template.jinja (the ATEM template)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, "79Labs/astraforge-8b-TCR")
messages = [{"role": "user", "content": "Price and latest news for Reliance?"}]
tools = [...] # OpenAI-style JSON schemas
ids = tok.apply_chat_template(messages, tools=tools, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=512, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=False))
| version | date | notes |
|---|---|---|
| 2.0.0 | 2026-10-05 | multi-call ATEM format, RAG, calculator reasoning; benchmarked against Claude on the same harness |
| 1.0.0 | earlier | tag v1.0.0: single-call agentic adapter with elicitation and confirm-first evaluations |