Downloads · 30 days
0
DAMO-NLP-SG/Video-LLaMA-Series
Video-LLaMA-Series is a visual question answering model from DAMO-NLP-SG. Use it for the visual question answering task on the model card, and read the license before you ship it in a product. The card lists the license as bsd-3-clause.
This is the Hugging Face repo for storing pre-trained & fine-tuned checkpoints of our Video-LLaMA, which is a multi-modal conversational large language model with video understanding capability.
Downloads · 30 days
0
Access
Public
Updated Jun 10, 2023
Repo size
2.7 GB
Likes
46
Public
Click a slice to open those files.
.pth2.7 GB · 100%
From the Hugging Face model README
This is the Hugging Face repo for storing pre-trained & fine-tuned checkpoints of our Video-LLaMA, which is a multi-modal conversational large language model with video understanding capability.
| Checkpoint | Link | Note |
|---|---|---|
| pretrain-vicuna7b | link | Pre-trained on WebVid (2.5M video-caption pairs) and LLaVA-CC3M (595k image-caption pairs) |
| finetune-vicuna7b-v2 | link | Fine-tuned on the instruction-tuning data from MiniGPT-4, LLaVA and VideoChat |
| pretrain-vicuna13b | link | Pre-trained on WebVid (2.5M video-caption pairs) and LLaVA-CC3M (595k image-caption pairs) |
| finetune-vicuna13b-v2 | link | Fine-tuned on the instruction-tuning data from MiniGPT-4, LLaVA and VideoChat |
| pretrain-ziya13b-zh | link | Pre-trained with Chinese LLM Ziya-13B |
| finetune-ziya13b-zh | link | Fine-tuned on machine-translated VideoChat instruction-following dataset (in Chinese) |
| pretrain-billa7b-zh | link | Pre-trained with Chinese LLM BiLLA-7B |
| finetune-billa7b-zh | link | Fine-tuned on machine-translated VideoChat instruction-following dataset (in Chinese) |
| Checkpoint | Link | Note |
|---|---|---|
| pretrain-vicuna7b | link | Pre-trained on WebVid (2.5M video-caption pairs) and LLaVA-CC3M (595k image-caption pairs) |
| finetune-vicuna7b-v2 | link | Fine-tuned on the instruction-tuning data from MiniGPT-4, LLaVA and VideoChat |
For launching the pre-trained Video-LLaMA on your own machine, please refer to our github repo.