Downloads · 30 days
14
16% of all-time downloads
InternRobotics/G2VLM-Qwen2-VL-2B
G2VLM-Qwen2-VL-2B is a image-text-to-text model from InternRobotics. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
<p align="left" <img src="https://huggingface.co/InternRobotics/G2VLM-2B-MoT/resolve/main/assets/icon.png" alt="G2VLM" width="200"/ </p
Downloads · 30 days
14
16% of all-time downloads
All-time downloads
87
Public
Parameters
2.4B
4.9 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors9.8 GB · 100%
From the Hugging Face model README
We present <b>G<sup>2</sup>VLM</b>, a geometry grounded vision-language model proficient in both spatial 3D reconstruction and spatial understanding tasks. For spatial reasoning questions, G<sup>2</sup>VLM can natively predict 3D geometry and employ interleaved reasoning for an answer.
This repository hosts the base model weights <b>BEFORE</b> the training of <b>G<sup>2</sup>VLM</b>, which is technically the same as Qwen2-VL-2B. Here we format it so it's easier for users to reproduce our trainings. For installation, usage instructions, and further documentation, please visit our GitHub repository.
<p align="left"><img src="https://huggingface.co/InternRobotics/G2VLM-2B-MoT/resolve/main/assets/teaser.png" width="100%"></p>G<sup>2</sup>VLM is a unified model that integrates both a geometric perception expert for 3D reconstruction and a semantic perception expert for multimodal understanding and spatial reasoning tasks. All tokens can do shared multi-modal self attention in each transformer block.
<p align="left"><img src="https://huggingface.co/InternRobotics/G2VLM-2B-MoT/resolve/main/assets/method.png" width="100%"></p>G2VLM is licensed under the Apache 2.0 license.
@article{hu2025g2vlmgeometrygroundedvision,
title={G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning},
author={Wenbo Hu and Jingli Lin and Yilin Long and Yunlong Ran and Lihan Jiang and Yifan Wang and Chenming Zhu and Runsen Xu and Tai Wang and Jiangmiao Pang},
year={2025},
eprint={2511.21688},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2511.21688},
}