Downloads · 30 days
8
7% of all-time downloads
fun-research/Video-LLaVA-Seg-Pretrain
Video-LLaVA-Seg-Pretrain is a machine learning model from fun-research. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
8
7% of all-time downloads
All-time downloads
119
Public
Parameters
8.7B
17.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors17.4 GB · 100%
How the weights are stored.
BF168.7B · 100%
From the Hugging Face model README
This is the official baseline implementation for the ViCas dataset. This is the pretrained model which has been optimized on a subset of WebVid10M and Panda70M for video captioning. The final model which has been finetuned on ViCaS is hosted here.
For details about setting up the model, refer to the Video-LLaVA-Seg GitHub repo.
For details about downloading and evaluating the dataset benchmark, refer to the ViCaS GitHub repo.