Downloads · 30 days
0
EPFL-VILAB/Video-4M-models
Video-4M-models is a image-to-video model from EPFL-VILAB. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as other.
Muhammad Uzair Khattak, Won Jun Kim, Reza Abbassi, Michael Murphy, Albias Havolli, Amir Zadeh, Chuan Li, Oğuzhan Fatih Kar, Muhammad Ferjad Naeem, Andrei Atanov, Roman Bachmann, Federico Tombari, Amir Zamir
Downloads · 30 days
0
Access
Public
Updated Sep 28, 2026
Repo size
20.1 GB
Likes
1
Public
Click a slice to open those files.
.pth13.4 GB · 100%
From the Hugging Face model README
Muhammad Uzair Khattak, Won Jun Kim, Reza Abbassi, Michael Murphy, Albias Havolli, Amir Zadeh, Chuan Li, Oğuzhan Fatih Kar, Muhammad Ferjad Naeem, Andrei Atanov, Roman Bachmann, Federico Tombari, Amir Zamir
EPFL, Google, Lambda
Official pre-trained checkpoints for Video-4M, an any-to-any multimodal model for the video domain that allows flexible traversal across both modality and time axes.
🌐 Project page | 📄 Paper | 💻 Code | 🤗 Tokenizers | 🤗 Demo | BibTeX
This repo holds the main Video-4M model checkpoints. We release two versions:
| Model | Params | File | Config |
|---|---|---|---|
| Video-4M-L | 705M | main_model/Video-4M-L.pth | Config |
| Video-4M-Pred-L | 705M | main_model/Video-4M-Pred-L.pth | Config |
Video-4M-Pred is an Video-4M model whose pretraining input-output mixture is biased toward predicting video from sparse modalities (e.g. caption, transcription) and toward forecasting tasks -- it can still perform all other any-to-any tasks.
For the tokenizers these checkpoints depend on, see EPFL-VILAB/Video-4M-tokenizers.
The easiest way to get started is the 🤗 Space demo, or the notebooks in the code repo, which download these checkpoints automatically:
from notebooks import pipeline
from notebooks.display_utils import video_grid
tokenizers = pipeline.load_tokenizers() # downloads the 7 tokenizers from the Hub on first call
model = pipeline.load_model(tokenizers, model_name="Video-4M_Pred_L") # downloads the model from the Hub on first call
example = pipeline.load_example(pipeline.list_examples()[0])
results = pipeline.generate(
example, input_modalities=["rgb"],
chain=pipeline.CHAIN_PRESETS["rgb_to_others"]["chain"],
tokenizers=tokenizers, model=model,
)
video_grid(results, cols=5, width=170, title="rgb -> everything else")
See the main repo README for installation, training, and evaluation instructions.
These model weights are released under the Sample Code license as found in LICENSE_WEIGHTS.
If you find this repository useful, please consider citing:
@article{video4m2026,
title={{Video-4M: Modeling the World Across Time and Modalities}},
author={Khattak, Muhammad Uzair and Kim, Won Jun and Abbassi, Reza and Havolli, Albias and Murphy, Michael and Zadeh, Amir and Li, Chuan and Kar, O\u{g}uzhan Fatih and Naeem, Muhammad Ferjad and Atanov, Andrei and Bachmann, Roman and Tombari, Federico and Zamir, Amir},
journal={arXiv preprint},
year={2026},
}