Downloads · 30 days
26
14% of all-time downloads
OpenMOSS-Team/Game-RL-LLaVA-OV-7B
Game-RL-LLaVA-OV-7B is a machine learning model from OpenMOSS-Team. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This model (GameQA-LLaVA-OV-7B) results from training LLaVA-OV-7B with GRPO solely on our GameQA-5K (sampled from the full GameQA-140K dataset).
Downloads · 30 days
26
14% of all-time downloads
All-time downloads
189
Public
Parameters
8B
16.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
This model (GameQA-LLaVA-OV-7B) results from training LLaVA-OV-7B with GRPO solely on our GameQA-5K (sampled from the full GameQA-140K dataset).
(The inference and evaluation configurations were unified across both the original open-source models and our trained models.)
This is the first work, to the best of our knowledge, that leverages game code to synthesize multimodal reasoning data for training VLMs. Furthermore, when trained with a GRPO strategy solely on GameQA (synthesized via our proposed Code2Logic approach), multiple cutting-edge open-source models exhibit significantly enhanced out-of-domain generalization.
[📖 Paper] [🤗 GameQA-140K Dataset] [🤗 GameQA-5K Dataset] [🤗 GameQA-InternVL3-8B ] [🤗 GameQA-Qwen2.5-VL-7B] [🤗 GameQA-LLaVA-OV-7B ]