Downloads · 30 days
0
m-Just/InSight-o3-vS
InSight-o3-vS is a image-text-to-text model from m-Just. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for adapter-transformers.
This is the vSearcher model introduced in paper "InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search". The model is finetuned from Qwen2.5-VL-7B-Instruct via RL as a subagent under vReas…
Downloads · 30 days
0
Access
Public
Updated Jan 29, 2026
Parameters
8.3B
33.2 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
This is the vSearcher model introduced in paper "InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search".
The model is finetuned from Qwen2.5-VL-7B-Instruct via RL as a subagent under vReasoner GPT-5-mini.
For more information on how to use this model, see our GitHub page.
@inproceedings{li2026insight_o3,
title={InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search},
author={Kaican Li and Lewei Yao and Jiannan Wu and Tiezheng Yu and Jierun Chen and Haoli Bai and Lu Hou and Lanqing Hong and Wei Zhang and Nevin L. Zhang},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026}
}