Downloads · 30 days
35K
4% of all-time downloads
google/vivit-b-16x2
vivit-b-16x2 is a video classification model from google. Use it for the video classification task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
ViViT model as introduced in the paper ViViT: A Video Vision Transformer by Arnab et al. and first released in this repository.
Downloads · 30 days
35K
4% of all-time downloads
All-time downloads
786K
Public
Repo size
1.1 GB
Likes
11
Public
Click a slice to open those files.
.bin356 MB · 100%
From the Hugging Face model README
ViViT model as introduced in the paper ViViT: A Video Vision Transformer by Arnab et al. and first released in this repository.
Disclaimer: The team releasing ViViT did not write a model card for this model so this model card has been written by the Hugging Face team.
ViViT is an extension of the Vision Transformer (ViT) to video.
We refer to the paper for details.
The model is mostly meant to intended to be fine-tuned on a downstream task, like video classification. See the model hub to look for fine-tuned versions on a task that interests you.
For code examples, we refer to the documentation.
@misc{arnab2021vivit,
title={ViViT: A Video Vision Transformer},
author={Anurag Arnab and Mostafa Dehghani and Georg Heigold and Chen Sun and Mario Lučić and Cordelia Schmid},
year={2021},
eprint={2103.15691},
archivePrefix={arXiv},
primaryClass={cs.CV}
}