Downloads · 30 days
86
9% of all-time downloads
Vision-CAIR/MiniGPT4-Video
MiniGPT4-Video is a machine learning model from Vision-CAIR. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
- Repository: https://github.com/Vision-CAIR/MiniGPT4-video - Paper: https://arxiv.org/abs/2407.12679
Downloads · 30 days
86
9% of all-time downloads
All-time downloads
995
Public
Parameters
7.8B
54.2 GB on disk
Likes
31
Public
Click a slice to open those files.
.safetensors9.7 GB · 67%
How the weights are stored.
I86.5B · 83%
From the Hugging Face model README
@misc{ataallah2024goldfishvisionlanguageunderstandingarbitrarily,
title={Goldfish: Vision-Language Understanding of Arbitrarily Long Videos},
author={Kirolos Ataallah and Xiaoqian Shen and Eslam Abdelrahman and Essam Sleiman and Mingchen Zhuge and Jian Ding and Deyao Zhu and Jürgen Schmidhuber and Mohamed Elhoseiny},
year={2024},
eprint={2407.12679},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2407.12679},
}
@misc{ataallah2024minigpt4videoadvancingmultimodalllms,
title={MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens},
author={Kirolos Ataallah and Xiaoqian Shen and Eslam Abdelrahman and Essam Sleiman and Deyao Zhu and Jian Ding and Mohamed Elhoseiny},
year={2024},
eprint={2404.03413},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2404.03413},
}