Downloads · 30 days
0
P4ddyki/MoTIF
MoTIF is a video classification model from P4ddyki. Use it for the video classification task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Read the Paper (arXiv) | GitHub Repository
Downloads · 30 days
0
Access
Public
Updated Feb 9, 2026
Repo size
3.5 GB
Likes
0
Public
Click a slice to open those files.
.pkl3.4 GB · 99%
From the Hugging Face model README
Read the Paper (arXiv) | GitHub Repository
Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but extending this paradigm to video remains challenging due to the difficulty of extracting concepts and modeling them over time. In this paper, we introduce MoTIF (Moving Temporal Interpretable Framework), a transformer-based concept architecture that operates on sequences of temporally grounded concept activations, by employing per-concept temporal self-attention to model when individual concepts recur and how their temporal patterns contribute to predictions. Central to the framework is an agentic concept discovery module to automatically extract object- and action-centric textual concepts from videos, yielding temporally expressive concept sets without manual supervision. Across multiple video benchmarks, this combination substantially narrows the performance gap between interpretable and black-box video models while maintaining faithful and temporally grounded concept explanations.
Create and activate an environment, then install requirements:
pip install -r requirements.txt
Place your datasets under Datasets/ (see the folder structure below). If you want to generate small demo clips or frames, you can use:
python save_videos.py
Compute (or recompute) the video/frame embeddings used by MoTIF:
python embedding.py
MoTIF’s training entry point is:
python train_MoTIF.py
Adjust hyperparameters in the script or via CLI flags (if exposed).
MoTIF.ipynb to visualize concept activations, attention over time, and example predictions.Models/ (see the notebook and code comments for expected paths).Pre-trained MoTIF checkpoints for all model variants are available on Hugging Face. The checkpoints include models trained on Breakfast, HMDB-51, and UCF-101 datasets with PE-L/14 backbone. We will upload soon additional checkpoints.
To use a pre-trained checkpoint, download it from the Hugging Face repository and place it in the Models/ directory. The notebook MoTIF.ipynb will automatically load the appropriate checkpoint based on the dataset and backbone you specify.
Please follow each dataset’s license and terms of use.
Note: If you use other datasets, you will need to adapt the dataset logic in the code (e.g., train/val/test splits, preprocessing, and loaders). Relevant places include utils/core/data/ (e.g., data.py, preprocessor.py, dataloader.py) and any dataset‑specific handling in embedding.py and train_MoTIF.py.
Datasets/ — dataset placeholdersEmbeddings/ — generated embeddings (created by scripts)Models/ — trained model checkpointsVideos/ — example videos used in the paper/one‑pagerutils/ — library code (vision encoder, projector, dataloaders, transforms, etc.)index.html — minimal one‑pager describing MoTIF (open locally in a browser)embedding.py, save_videos.py, train_MoTIF.py — main scriptsMoTIF.ipynb — notebook for inspection and visualizationIf you use MoTIF in your research, please consider citing:
@misc{knab2025conceptsmotiontemporalbottlenecks,
title={Concepts in Motion: Temporal Bottlenecks for Interpretable Video Classification},
author={Patrick Knab and Sascha Marton and Philipp J. Schubert and Drago Guggiana and Christian Bartelt},
year={2025},
eprint={2509.20899},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2509.20899},
}
utils/core codebase are adapted from the Perception Encoder framework.For questions and discussion, please open an issue or contact the authors.