Downloads · 30 days
29
6% of all-time downloads
TIGER-Lab/VL-Reasoner-72B
VL-Reasoner-72B is a visual question answering model from TIGER-Lab. Use it for the visual question answering task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
🚀 News: <uWe release our meticulously curated collection of RL training queries for multimodal reasoning: ViRL39K.</u
Downloads · 30 days
29
6% of all-time downloads
All-time downloads
455
Public
Parameters
73.4B
147 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors147 GB · 100%
From the Hugging Face model README
🚀 News: <u>We release our meticulously curated collection of RL training queries for multimodal reasoning: ViRL39K.</u>
VL-Reasoner-72B achieves superior results on various multimodal reasoning benchmarks.
It is trained using the GRPO-SSR techniques, serving as the foundation for VL-Rethinker.
For details of our approach and performance comparison, please see our paper.
For details of training and evaluation, please see our code repo.
Explore further via the following links:
| 🚀Project Page | 📖Paper | 🔗Github | 🤗Data |
If you feel this model useful, please give us a free cite:
@article{vl-rethinker,
title={VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning},
author = {Wang, Haozhe and Qu, Chao and Huang, Zuming and Chu, Wei and Lin, Fangzhen and Chen, Wenhu},
journal={arXiv preprint arXiv:2504.08837},
year={2025}
}