Downloads · 30 days
8
0% of all-time downloads
Kangheng/OVR-7B-ColdStart
OVR-7B-ColdStart is a machine learning model from Kangheng. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
8
0% of all-time downloads
All-time downloads
6.5K
Public
Parameters
8.3B
16.6 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
How the weights are stored.
BF168.3B · 100%
From the Hugging Face model README


The remarkable reasoning capbility of Large Language Models (LLMs) stems from cognitive behaviors that emerge when reinforcing against verifiable rewards. This work investigates how to transfer this principle to Multimodal LLMs (MLLMs) to unlock advanced visual reasoning.
We introduce a two-stage paradigm built on Qwen2.5-VL-7B: a massive text-only cold-start fine-tuning, followed by multimodal reinforcement learning (RL) spanning nearly 1,000 steps—surpassing all prior open-source efforts in scale. This pioneering work reveals three fundamental insights:
Our resulting model, Open-Vision-Reasoner (OVR), achieves state-of-the-art performance on a suite of reasoning benchmarks, including 95.3% on MATH500, 51.8% on MathVision and 54.6% on MathVerse. We release our model, data, and training dynamics to catalyze the development of more capable, behavior-aligned multimodal reasoners.
| Model | Description | Download |
|---|---|---|
| OVR-7B-ColdStart | Intermediate model after massive language-only cold-start fine-tuning | 🤗 OVR-7B-ColdStart |
| OVR-7B-RL | Final model after large-scale multimodal RL training | 🤗 OVR-7B-RL |

vllm serve Kangheng/OVR-7B-ColdStart --port 8000 --host 0.0.0.0 --tensor-parallel-size 1 --gpu-memory-utilization 0.6