Downloads · 30 days
18
14% of all-time downloads
chamber111/VPPO-32B
VPPO-32B is a machine learning model from chamber111. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
VPPO-32B is a state-of-the-art Large Vision-Language Model (LVLM) specialized for complex multimodal reasoning tasks. It is the 32B parameter version of our model, fine-tuned from Qwen2.5-VL-32B-Instruct using a novel…
Downloads · 30 days
18
14% of all-time downloads
All-time downloads
125
Public
Parameters
33.5B
66.9 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors66.9 GB · 100%
From the Hugging Face model README
VPPO-32B is a state-of-the-art Large Vision-Language Model (LVLM) specialized for complex multimodal reasoning tasks. It is the 32B parameter version of our model, fine-tuned from Qwen2.5-VL-32B-Instruct using a novel reinforcement learning algorithm called Visually-Perceptive Policy Optimization (VPPO).
The core innovation of VPPO is its ability to solve the "uniform learning signal" problem that plagues standard RL fine-tuning. Instead of broadcasting a single reward to all tokens in a reasoning chain, VPPO intelligently identifies and focuses policy updates on the sparse, critical tokens that are highly dependent on visual input. This hierarchical "spotlight" mechanism allows the model to develop a more robust and genuine perception-grounded reasoning capability.
As a result, VPPO-32B demonstrates significant performance improvements over strong baselines across a wide range of challenging benchmarks, including mathematics, geometry, and logic problems. It also exhibits superior training stability and faster convergence.
Qwen/Qwen2.5-VL-32B-InstructVPPO-RL2510.09285The model was fine-tuned on ViRL39K, a diverse collection of multimodal reasoning problems. The original dataset can be found on the Hugging Face Hub: TIGER-Lab/ViRL39K.
The model was trained using our Visually-Perceptive Policy Optimization (VPPO) algorithm, which is a modification of the Group Relative Policy Optimization (GRPO) framework. The procedure involves generating responses, calculating token-level visual dependency, and using this dependency to shape the advantage and filter gradients during the policy update step.
The model was evaluated on a comprehensive suite of 8 diverse multimodal reasoning benchmarks:
Performance is measured by average accuracy@8, which is the average success rate over 8 independent generations per problem (at temperature=1.0) using exact-match scoring.
If you use this model in your work, please cite our paper:
BibTeX:
@article{huang2025spotlight,
title={Spotlight on Token Perception for Multimodal Reinforcement Learning},
author={Huang, Siyuan and Qu, Xiaoye and Li, Yafu and Luo, Yun and He, Zefeng and Liu, Daizong and Cheng, Yu},
journal={arXiv preprint arXiv:2510.09285},
year={2025}
}