Downloads · 30 days
2.4K
43% of all-time downloads
allenai/tmax-2b
tmax-2b is a machine learning model from allenai. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<p align="center" 💻 <a href="https://github.com/hamishivi/tmax"Code</a · 🤗 <a href="https://huggingface.co/collections/allenai/tmax"Models & Data</a · 📜 <a href="https://arxiv.org/abs/2606.23321"Paper</a ·
Downloads · 30 days
2.4K
43% of all-time downloads
All-time downloads
5.6K
Public
Parameters
1.9B
11.3 GB on disk
Likes
4
Public
Click a slice to open those files.
.safetensors3.8 GB · 99%
From the Hugging Face model README

[!NOTE] For full information, go check out the Tmax paper here.
TMax 2B is a model trained using DPPO on top of Qwen 3.5 2B for use as a terminal-agent. It achieves roughly 4% on Terminal Bench 2.0 after 100 steps of RL training.

This model is part of a collection of terminal agents in various sizes.
Additionally, we provide model checkpoints as branches of the repository. The main model checkpoint is step 100 as this performed best on TBLite.
| Model | TB Lite | TB 2.1 | TB 2.0 (daytona) |
|---|---|---|---|
| Qwen 3.5 2B | 5.71 +/- 1.6 | 1.9 +/- 1.4 | 2.3 +/- 1.0 |
| Tmax 2B | 11.8 +/- 1.4 | 4.2 +/- 1.2 | 2.9 +/- 0.6 |
| Qwen 3.5 4B | 31.8 +/- 3.8 | ? | 16.6 +/- 1.7 |
| Tmax 4B (this model!) | 42.6 +/- 1.5 | 19.9 +/- 1.1 | 18.9 +/- 1.9 |
| Qwen 3.5 9B | 41.9 +/- 2.7 | 16.1 +/- 3.7 | 21.1 +/- 2.6 |
| Tmax 9B | 57.2 +/- 2.5 | 28.8 +/- 3.7 | 27.2 +/- 1.5 |
| Qwen 3.6 27B | 70.8 +/- 2.1 | 40.5 +/- 2.4 | 39.6 +/- 2.1 |
| Tmax 27B | 68.6 +/- 4.7 | 44.9 +/- 1.8 | 42.7 +/- 0.7 |
For details on evaluation methodology please check our paper. In general, we used a podman (docker) backend with default timeouts and custom harness similar to mini-swe-agent. For the 'daytona' runs, we used the daytona backend. For Lite/2.1, we show mean and standard error over 3 runs. For daytona, we show it over 5 runs.
To use this model, we recommend serving with vllm (or your inference framework of choice) with:
uvx vllm==0.19.1 serve allenai/tmax-2b \
--served-model-name tmax-2b \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--port 8008 \
--max-model-len 65536 \
--tensor-parallel-size 8 \
--language_model_only
Make sure to set language_model_only as we removed the vision head during training.
For more details on evaluation, please see our codebase.
This model was trained using DPPO with the following hyperparameters:
For more details on training, please see our codebase.
This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
If you use our model or data, please cite our paper:
@misc{ivison2026tmaxsimplerecipeterminal,
title={Tmax: A simple recipe for terminal agents},
author={Hamish Ivison and Junjie Oscar Yin and Rulin Shao and Teng Xiao and Nathan Lambert and Hannaneh Hajishirzi},
year={2026},
eprint={2606.23321},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2606.23321},
}