Downloads · 30 days
82
16% of all-time downloads
mPLUG/ToolCUA-8B
ToolCUA-8B is a image-text-to-text model from mPLUG. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers/. The card lists the license as mit.
<h1 style=" font-family:-apple-system,BlinkMacSystemFont,'Segoe UI',Helvetica,Arial,sans-serif; font-size:48px; font-weight:700; line-height:1.25; text-align:center; margin:0 0 24px;" ToolCUA-8B </h1
Downloads · 30 days
82
16% of all-time downloads
All-time downloads
497
Public
Parameters
8.8B
35.1 GB on disk
Likes
6
Public
Click a slice to open those files.
.safetensors35.1 GB · 100%
From the Hugging Face model README
<a href="https://x-plug.github.io/ToolCUA/" style=" display:inline-block; padding:8px 24px; background:#2b2b2b; color:#ffffff; border-radius:36px; text-decoration:none; font-weight:600; font-size:16px;"> 🌐 Website </a>
<a href="https://arxiv.org/abs/2605.12481" style=" display:inline-block; padding:8px 24px; background:#2b2b2b; color:#ffffff; border-radius:36px; text-decoration:none; font-weight:600; font-size:16px;"> 📑 Paper </a>
<a href="https://github.com/X-PLUG/ToolCUA" style=" display:inline-block; padding:8px 24px; background:#2b2b2b; color:#ffffff; border-radius:36px; text-decoration:none; font-weight:600; font-size:16px;"> 💻 Code </a>
</div>ToolCUA-8B is an end-to-end computer-use agent for orchestrating GUI actions and structured tool calls. It learns when to continue through GUI interaction, when to invoke tools, and when to switch back, enabling shorter and more reliable desktop task trajectories.
<p align="center"> <img src="https://github.com/X-PLUG/ToolCUA/raw/main/assets/main_teaser.png" width="760" alt="ToolCUA teaser"> </p>ToolCUA uses a staged training pipeline for GUI-Tool path selection:
On feasible OSWorld-MCP tasks, ToolCUA-8B reaches 46.85% overall accuracy, 24.32% Tool Invocation Rate (TIR), and 14.93 Average Completion Steps (ACS). Compared with Qwen3-VL-8B-Instruct, it improves accuracy by +18.62, improves TIR by +15.91, and reduces ACS by 4.41 steps.
<p align="center"> <img src="https://github.com/X-PLUG/ToolCUA/raw/main/assets/main_results.png" width="760" alt="ToolCUA main results"> </p> <p align="center"> <img src="https://github.com/X-PLUG/ToolCUA/raw/main/assets/app_results.png" width="760" alt="ToolCUA application results"> </p>We recommend vLLM for deployment. Use vllm>=0.12.0 and enable --trust-remote-code.
MAX_IMAGE=${MAX_IMAGE:-5}
IMAGE_LIMIT_ARGS='{"image": '"$MAX_IMAGE"'}'
PIXEL_ARGS='{"size": {"longest_edge": 3072000, "shortest_edge": 65536}}'
vllm serve X-PLUG/ToolCUA-8B \
--trust-remote-code \
--max-model-len 32768 \
--mm-processor-kwargs "$PIXEL_ARGS" \
--limit-mm-per-prompt "$IMAGE_LIMIT_ARGS" \
--tensor-parallel-size 1 \
--allowed-local-media-path '/' \
--port 4243 \
--gpu-memory-utilization 0.85 \
--mm-processor-cache-gb 0 \
--no-enable-prefix-caching \
--enforce-eager \
--max-logprobs 50
@article{hu2026toolcua,
title={ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents},
author={Hu, Xuhao and Zhang, Xi and Xu, Haiyang and Qiao, Kyle and Yang, Jingyi and Huang, Xuanjing and Shao, Jing and Yan, Ming and Ye, Jieping},
journal={arXiv preprint arXiv:2605.12481},
year={2026}
}