Downloads · 30 days
75
2% of all-time downloads
neulab/UIX-Qwen2
UIX-Qwen2 is a machine learning model from neulab. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as odc-by.
Downloads · 30 days
75
2% of all-time downloads
All-time downloads
3.2K
Public
Parameters
8B
16.1 GB on disk
Likes
22
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
🌐 Homepage | 🐍 GitHub | 📖 arXiv
We introduce MultiUI, a dataset containing 7.3 million samples from 1 million websites, covering diverse multi- modal tasks and UI layouts. Models trained on MultiUI not only excel in web UI tasks—achieving up to a 48% improvement on VisualWebBench and a 19.1% boost in action accuracy on a web agent dataset Mind2Web—but also generalize surprisingly well to non-web UI tasks and even to non-UI domains, such as document understanding, OCR, and chart interpretation.
<video controls autoplay src="https://cdn-uploads.huggingface.co/production/uploads/65403d8781a8731a1c09a584/vk7yT4Y7ydBOHM6BojmlI.mp4"></video>
The model training is based on the LLaVA-NeXT.
For deployment, refer to SGLang deployment section in LLaVA-NeXT repo.
For benchmark evaluation, the awesome lmms-eval package is used. Check our repo MultiUI to evaluate on benchmarks mentioned in the paper.



If you find this work helpful, please cite out paper:
@misc{liu2024harnessingwebpageuistextrich,
title={Harnessing Webpage UIs for Text-Rich Visual Understanding},
author={Junpeng Liu and Tianyue Ou and Yifan Song and Yuxiao Qu and Wai Lam and Chenyan Xiong and Wenhu Chen and Graham Neubig and Xiang Yue},
year={2024},
eprint={2410.13824},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2410.13824},
}