Downloads · 30 days
142
0% of all-time downloads
inclusionAI/ViLaSR
ViLaSR is a image-text-to-text model from inclusionAI. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product.
This repository contains the ViLaSR-7B model as presented in Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing.
Downloads · 30 days
142
0% of all-time downloads
All-time downloads
30.1K
Public
Parameters
8.3B
16.6 GB on disk
Likes
18
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
This repository contains the ViLaSR-7B model as presented in Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing.
Please refer to the code https://github.com/AntResearchNLP/ViLaSR.
@misc{wu2025reinforcingspatialreasoningvisionlanguage,
title={Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing},
author={Junfei Wu and Jian Guan and Kaituo Feng and Qiang Liu and Shu Wu and Liang Wang and Wei Wu and Tieniu Tan},
year={2025},
eprint={2506.09965},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2506.09965},
}