Downloads · 30 days
0
cythu/PeBR_R1
PeBR_R1 is a machine learning model from cythu. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Oct 17, 2025
Repo size
24.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors24.7 GB · 100%
From the Hugging Face model README
Yan Chen<sup>1</sup>, Long Li<sup>2</sup>, Teng Xi<sup>2</sup>, Long Zeng<sup>1✉️</sup>, Jingdong Wang<sup>2</sup>
<br/><sup>1</sup>Tsinghua University, <sup>2</sup>Baidu
<sup>✉️</sup>Corresponding Author
</div>Reinforcement learning (RL) has proven highly effective in eliciting the reasoning capabilities of large language models (LLMs). Inspired by this success, recent studies have explored applying similar techniques to vision-language models (VLMs), aiming to enhance their reasoning performance.
However, directly transplanting RL methods from LLMs to VLMs is suboptimal, as the tasks faced by VLMs are inherently more complex. Specifically, VLMs must first accurately perceive and understand visual inputs before reasoning can be effectively performed.
To address this challenge, we propose a two-stage reinforcement learning framework designed to jointly enhance both the perceptual and reasoning capabilities of VLMs. To mitigate the vanishing advantage issue commonly observed in RL training, we first perform dataset-level sampling to selectively strengthen specific capabilities using distinct data sources.
During training:
After the proposed two-stage reinforcement learning process, we obtain PeBR-R1, a vision-language model with significantly enhanced perceptual and reasoning capabilities.
Experimental results on seven benchmark datasets demonstrate the effectiveness of our approach and validate the superior performance of PeBR-R1 across diverse visual reasoning tasks.
If you find PeBR-R1 useful for your research, please consider citing our work:
@article{chen2025perception,
title={Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models},
author={Chen, Yan and Li, Long and Xi, Teng and Zeng, Long and Wang, Jingdong},
journal={arXiv preprint arXiv:2509.13031},
year={2025}
}