Downloads · 30 days
0
katefgroup/UniVLG
UniVLG is a image-text-to-text model from katefgroup. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-4.0.
This repository contains the UniVLG model, as presented in Unifying 2D and 3D Vision-Language Understanding. UniVLG is a unified architecture for 2D and 3D vision-language understanding.
Downloads · 30 days
0
Access
Public
Updated May 2, 2025
Repo size
12.4 GB
Likes
0
Public
Click a slice to open those files.
.pth7.4 GB · 100%
From the Hugging Face model README
This repository contains the UniVLG model, as presented in Unifying 2D and 3D Vision-Language Understanding. UniVLG is a unified architecture for 2D and 3D vision-language understanding.
Project page: https://univlg.github.io
The model uses a custom loading tool (uvx). Checkpoints are available on Hugging Face: Hugging Face. See the GitHub repository for code and instructions.
@article{jain2025unifying,
title={Unifying 2D and 3D Vision-Language Understanding},
author={Jain, Ayush and Swerdlow, Alexander and Wang, Yuzhou and Arnaud, Sergio and Martin, Ada and Sax, Alexander and Meier, Franziska and Fragkiadaki, Katerina},
journal={arXiv preprint arXiv:2503.10745},
year={2025}
}
License Note: The majority of UniVLG is licensed under CC-BY-NC, however, portions of the project (specifically Odin and Pointcept) are available under separate MIT license terms. Please refer to the GitHub repository for details.