Downloads ยท 30 days
27
16% of all-time downloads
GAIR/daVinci-Agency
daVinci-Agency is a text generation model from GAIR. Use it when you need the model to write or continue text. The card lists the license as mit.
<div style="display: flex; justify-content: center; align-items: center; gap: 20px;" <img src="assets/sii.jpg" alt="SII" width="100px" <img src="assets/asi.png" alt="ASI" width="100px" </div
Downloads ยท 30 days
27
16% of all-time downloads
All-time downloads
168
Public
Parameters
353B
706 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors706 GB ยท 100%
How the weights are stored.
BF16353B ยท 100%
From the Hugging Face model README
daVinci-Agency is a long-horizon agentic model finetuned on GLM-4.6. It is designed to master long-horizon tasks by learning from the iterative evolutionary process of real-world software development.
Unlike standard instruction-tuned models, daVinci-Agency is trained on trajectories that explicitly embody key meta-skills: task decomposition, long-term consistency, and iterative refinement. These trajectories, mined from real GitHub Pull Request chains, average 85k tokens and 116 tool invocations, enabling the model to handle complex, multi-step agentic workflows with superior proficiency.
daVinci-Agency (239 samples) consistently outperforms much larger synthetic datasets.
| Training Data | Samples | SWE-bench(SWE-agent) | Toolathlon | $\tau^2$-bench | Overall Avg. |
|---|---|---|---|---|---|
| GLM-4.6 (Base) | - | 0.608 | 0.157 | 0.675 | 0.441 |
| SWE-Smith | 66,000 | 0.404 | 0.093 | 0.586 | 0.373 |
| CC-bench | 260 | 0.618 | 0.000 | 0.697 | 0.436 |
| daVinci-Agency | 239 | 0.632 | 0.231 | 0.707 | 0.475 |
Baseline Models Comparison
| Model | SWE-bench(SWE-agent) | Toolathlon | AgencyBench | Overall<br>Avg. |
|---|---|---|---|---|
| DeepSeek-v3.2 | 0.456 | 0.250 | 11.6 | 0.366 |
| Qwen3-235B | 0.504 | 0.046 | 4.6 | 0.309 |
| Kimi-K2-Thinking | 0.318 | 0.213 | 11.8 | 0.404 |
| GLM-4.6 | 0.608 | 0.157 | 11.9 | 0.441 |
| GLM-4.6-daVinci-Agency | 0.632 | 0.231 | 15.9 | 0.475 |
Since daVinci-Agency is based on GLM-4.6, it utilizes the standard GLM tokenizer and chat template.
Following the GLM-4.6 guidelines, we recommend the following parameters for optimal performance, especially in code-related tasks:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
device = "cuda"
model_path = "GAIR/daVinci-Agency" # Replace with actual path
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype=torch.bfloat16,
low_cpu_mem_usage=True,
trust_remote_code=True
).to(device).eval()
# Example: Complex Long-Horizon Task
query = "Refactor the authentication module in this repository to support OAuth2, ensuring backward compatibility."
messages = [
{"role": "system", "content": "You are an intelligent software engineering agent capable of long-horizon planning and execution."},
{"role": "user", "content": query}
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_tensors="pt",
return_dict=True
).to(device)
gen_kwargs = {
"max_new_tokens": 4096,
"do_sample": True,
"top_p": 0.95,
"temperature": 1.0,
"top_k": 40
}
with torch.no_grad():
outputs = model.generate(**inputs, **gen_kwargs)
response = outputs[:, inputs['input_ids'].shape[1]:]
print(tokenizer.decode(response[0], skip_special_tokens=True))
If you use daVinci-Agency in your research, please cite our work:
@article{jiang2026davinci,
title={daVinci-Agency: Unlocking Long-Horizon Agency Data-Efficiently},
author={Mohan Jiang and Dayuan Fu and Junhao Shi and Ji Zeng and Weiye Si and Keyu Li and Xuefeng Li and Yang Xiao and Wenjie Li and Dequan Wang and Pengfei Liu},
journal={arXiv preprint arXiv:2602.02619},
year={2026}
}