Downloads · 30 days
0
ZhenyangLiu/ActiveVLA
ActiveVLA is a machine learning model from ZhenyangLiu. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
0
Access
Public
Updated Jan 16, 2026
Repo size
—
Likes
2
Public
Click a slice to open those files.
.md4.2 KB · 74%
From the Hugging Face model README
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
Zhenyang Liu<sup>1,2</sup>, Yongchong Gu<sup>1</sup>, Yikai Wang<sup>3</sup>,
Xiangyang Xue<sup>1,†</sup>, Yanwei Fu<sup>1,2,†</sup>
<sup>1</sup>Fudan University, <sup>2</sup>Shanghai Innovation Institute, <sup>3</sup>Nanyang Technological University
<sup>†</sup>Corresponding Authors
</div>This repository is the official implementation of ActiveVLA. We are currently preparing the code and data for release. Please stay tuned!
Most existing Vision-Language-Action (VLA) models rely on static, wrist-mounted cameras that provide a fixed, end-effector-centric viewpoint. This setup limits perceptual flexibility: the agent cannot adaptively adjust its viewpoint or camera resolution according to the task context, leading to failures in long-horizon tasks or fine-grained manipulation due to occlusion and lack of detail.
We propose ActiveVLA, a novel vision-language-action framework that explicitly integrates active perception into robotic manipulation. Unlike passive perception methods, ActiveVLA empowers robots to:
By dynamically refining its perceptual input, ActiveVLA achieves superior adaptability and performance in complex scenarios. Experiments show that ActiveVLA outperforms state-of-the-art baselines on RLBench, COLOSSEUM, and GemBench, and transfers seamlessly to real-world robots.
We propose a coarse-to-fine active perception framework that integrates 3D spatial reasoning with vision-language understanding.
The pipeline consists of two main stages:
Note: For more visualizations and real-world robot demos, please visit our Project Page.
ActiveVLA achieves state-of-the-art performance across multiple benchmarks:
If you find our work useful in your research, please consider citing:
@misc{liu2026activevlainjectingactiveperception,
title={ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation},
author={Zhenyang Liu and Yongchong Gu and Yikai Wang and Xiangyang Xue and Yanwei Fu},
year={2026},
eprint={2601.08325},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2601.08325},
}