Downloads · 30 days
10.6K
39% of all-time downloads
mPLUG/GUI-Owl-1.5-2B-Instruct
GUI-Owl-1.5-2B-Instruct is a machine learning model from mPLUG. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
<img src="https://github.com/X-PLUG/MobileAgent/blob/main/Mobile-Agent-v3.5/assets/guiowl15logo.png?raw=true" width="80%"/
Downloads · 30 days
10.6K
39% of all-time downloads
All-time downloads
27.4K
Public
Parameters
2.1B
4.3 GB on disk
Likes
23
Trending 1
Click a slice to open those files.
.safetensors4.3 GB · 100%
From the Hugging Face model README
GUI-Owl 1.5 is the next-generation native GUI agent model family built on Qwen3-VL. It supports multi-platform GUI automation across desktops, mobile devices, browsers, and more. Powered by a scalable hybrid data flywheel, unified agent capability enhancement, and multi-platform environment RL (MRPO), GUI-Owl 1.5 offers a full spectrum of models.
Key highlights:
| Model | OSWorld-Verified | AndroidWorld | OSWorld-MCP | Mobile-World | WindowsAA | WebArena | VisualWebArena | WebVoyager | Online-Mind2Web |
|---|---|---|---|---|---|---|---|---|---|
| GUI-Owl-1.5-2B-Instruct | 43.5 | 67.9 | 33.0 | 31.3 | 25.8 | - | - | - | - |
| GUI-Owl-1.5-4B-Instruct | 48.2 | 69.8 | 31.7 | 32.3 | 29.4 | - | - | - | - |
| GUI-Owl-1.5-8B-Instruct | 52.3 | 69.0 | 41.8 | 41.8 | 31.7 | 45.7 | 39.4 | 69.9 | 41.7 |
| GUI-Owl-1.5-8B-Thinking | 52.9 | 71.6 | 38.8 | 33.3 | 35.1 | 46.7 | 40.8 | 78.1 | 48.6 |
| GUI-Owl-1.5-32B-Instruct | 56.5 | 69.4 | 47.6 | 46.8 | 44.8 | - | - | - | - |
| GUI-Owl-1.5-32B-Thinking | 56.0 | 68.2 | 43.8 | 42.8 | 44.1 | 48.4 | 46.6 | 82.1 | - |
Please refer to the technical report for detailed results on ScreenSpot-v2, ScreenSpot-Pro, OSWorld-G, MMBench-GUI, and more.
Please refer to our cookbook.
We recommand deploy GUI-Owl-1.5 through vllm
This script has been validated on an A100 with 96 GB of VRAM.
PIXEL_ARGS='{"size": {"longest_edge": 3072000, "shortest_edge": 65536}}'
IMAGE_LIMIT_ARGS='image=5'
MP_SIZE=1
vllm serve $CKPT \
--max-model-len 32768 \
--mm-processor-kwargs "$PIXEL_ARGS" \
--limit-mm-per-prompt "$IMAGE_LIMIT_ARGS" \
--tensor-parallel-size $MP_SIZE \
--allowed-local-media-path '/' \
--port 4243 \
If you find this model useful, please cite our paper:
@article{MobileAgentv3.5,
title={Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents},
author={Haiyang Xu, Xi Zhang, Haowei Liu, Junyang Wang, Zhaozai Zhu, Shengjie Zhou, Xuhao Hu, Feiyu Gao, Junjie Cao, Zihua Wang, Zhiyuan Chen, Jitong Liao, Qi Zheng, Jiahui Zeng, Ze Xu, Shuai Bai, Junyang Lin, Jingren Zhou, Ming Yan},
journal={arXiv preprint arXiv:2602.16855},
year={2026}
}