Downloads · 30 days
0
innat/videoswin
videoswin is a video classification model from innat. Use it for the video classification task on the model card, and read the license before you ship it in a product. It is set up for tf-keras. The card lists the license as mit.
Downloads · 30 days
0
Access
Public
Updated Jul 6, 2024
Repo size
2.6 GB
Likes
5
Public
Click a slice to open those files.
.data-00000-of-000011.4 GB · 53%
From the Hugging Face model README

VideoSwin is a pure transformer based video modeling algorithm, attained top accuracy on the major video recognition benchmarks. In this model, the author advocates an inductive bias of locality in video transformers, which leads to a better speed-accuracy trade-off compared to previous approaches which compute self-attention globally even with spatial-temporal factorization. The locality of the proposed video architecture is realized by adapting the Swin Transformer designed for the image domain, while continuing to leverage the power of pre-trained image models.
This is a unofficial Keras implementation of Video Swin transformers. The official PyTorch implementation is here based on mmaction2.
The 3D swin-video checkpoints are listed in MODEL_ZOO.md. Following are some hightlights.
In the training phase, the video swin mdoels are initialized with the pretrained weights of image swin models. In that case, IN referes to ImageNet.
| Backbone | Pretrain | Top-1 | Top-5 | #params | FLOPs | config |
|---|---|---|---|---|---|---|
| Swin-T | IN-1K | 78.8 | 93.6 | 28M | ? | swin-t |
| Swin-S | IN-1K | 80.6 | 94.5 | 50M | ? | swin-s |
| Swin-B | IN-1K | 80.6 | 94.6 | 88M | ? | swin-b |
| Swin-B | IN-22K | 82.7 | 95.5 | 88M | ? | swin-b |
| Backbone | Pretrain | Top-1 | Top-5 | #params | FLOPs | config |
|---|---|---|---|---|---|---|
| Swin-B | IN-22K | 84.0 | 96.5 | 88M | ? | swin-b |
| Backbone | Pretrain | Top-1 | Top-5 | #params | FLOPs | config |
|---|---|---|---|---|---|---|
| Swin-B | Kinetics 400 | 69.6 | 92.7 | 89M | ? | swin-b |