Downloads · 30 days
14
7% of all-time downloads
godx7/PixelEyes-4B
PixelEyes-4B is a image-text-to-text model from godx7. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product.
<p align="center" <h1 align="center"👀 PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking</h1 </p
Downloads · 30 days
14
7% of all-time downloads
All-time downloads
200
Public
Parameters
4.8B
9.7 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors9.7 GB · 100%
From the Hugging Face model README
PixelEyes enhances active visual search in MLLMs by delegating fine-grained localization to a specialized perception tool, thereby achieving efficient and accurate multi-turn visual reasoning.
This repository contains the weights for PixelEyes, a multi-turn visual reasoning agent that explicitly decouples reasoning from perception, introduced in the paper PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking.
For more details on environment setup, training, and evaluation, please visit the GitHub repository.
If you find this project helpful in your research, please cite our paper:
@misc{gong2026pixeleyesdecouplingperceptionreasoning,
title={PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking},
author={Dengxian Gong and Yuanzheng Wu and Haobo Yuan and Zhengdong Hu and Tao Zhang and Yikang Zhou and Shihao Chen and Quanzhu Niu and Kai Wang and Jason Li and Haochen Wang and Lu Qi and Shunping Ji and Ming-Hsuan Yang},
year={2026},
eprint={2607.00115},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2607.00115},
}