Downloads · 30 days
52
8% of all-time downloads
DAMO-NLP-SG/VideoRefer-VideoLLaMA3-2B
VideoRefer-VideoLLaMA3-2B is a video-text-to-text model from DAMO-NLP-SG. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
<p align="center" <img src="https://cdn-uploads.huggingface.co/production/uploads/64a3fe3dde901eb01df12398/ZrZPYT0Q3wgza7Vc5BmyD.png" width="100%" style="margin-bottom: 0.2;"/ <p
Downloads · 30 days
52
8% of all-time downloads
All-time downloads
658
Public
Parameters
2B
3.9 GB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors3.9 GB · 100%
From the Hugging Face model README
If you find VideoRefer Suite useful for your research and applications, please cite using this BibTeX:
@InProceedings{Yuan_2025_CVPR,
author = {Yuan, Yuqian and Zhang, Hang and Li, Wentong and Cheng, Zesen and Zhang, Boqiang and Li, Long and Li, Xin and Zhao, Deli and Zhang, Wenqiao and Zhuang, Yueting and Zhu, Jianke and Bing, Lidong},
title = {VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM},
booktitle = {Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)},
month = {June},
year = {2025},
pages = {18970-18980}
}
@article{damonlpsg2025videollama3,
title={VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding},
author={Boqiang Zhang, Kehan Li, Zesen Cheng, Zhiqiang Hu, Yuqian Yuan, Guanzheng Chen, Sicong Leng, Yuming Jiang, Hang Zhang, Xin Li, Peng Jin, Wenqi Zhang, Fan Wang, Lidong Bing, Deli Zhao},
journal={arXiv preprint arXiv:2501.13106},
year={2025},
url = {https://arxiv.org/abs/2501.13106}
}