Downloads Β· 30 days
0
EPFL-VILAB/Video-4M-tokenizers
Video-4M-tokenizers is a machine learning model from EPFL-VILAB. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
π Project page | π Paper | π» Code | π€ Models | π€ Demo
Downloads Β· 30 days
0
Access
Public
Updated Oct 5, 2026
Repo size
9.6 GB
Likes
1
Public
Click a slice to open those files.
.ckpt9.6 GB Β· 100%
From the Hugging Face model README
π Project page | π Paper | π» Code | π€ Models | π€ Demo
This repository hosts the modality tokenizers used by Video-4M, where each modality is represented as a sequence of discrete tokens using its own dedicated tokenizer.
Let's briefly go through the tokenization details for each modality:
Video-4M operates on 4-second clips (17 frames sampled at 4 FPS) at 128x128 resolution. For the feature map modalities, features are extracted at each extractor's default resolution before tokenization, i.e., 224x224 for DINOv2 and SigLIP 2, and 256x256 for V-JEPA 2.
| Modality | Path in repo | Tokens per clip | Vocab size |
|---|---|---|---|
| RGB | rgb/ckpt.ckpt | 1280 | 32,768 |
| Depth | depth/ckpt.ckpt | 1280 | 32,768 |
| Surface normals | surface-normals/ckpt.ckpt | 1280 | 32,768 |
| Optical flow | opticalflow/ckpt.ckpt | 1280 | 32,768 |
| V-JEPA 2 | v-jepa-2/ckpt.ckpt | 1024 | 16,807 |
| DINOv2 | dinov2/ckpt.ckpt | 1280 | 16,807 |
| SigLIP 2 | siglip-2/ckpt.ckpt | 980 | 16,807 |
The remaining modalities are not hosted here:
EPFL-VILAB/4M_tokenizers_human-poses_1k_8.The notebooks in our main GitHub code repository download these tokenizers automatically. Below, we provide an example of loading the RGB tokenizer directly and using its encode/decode calls:
import torch
from notebooks import pipeline
tokenizers = pipeline.load_tokenizers() # downloads the tokenizers from the Hub on first call
rgb_tokenizer = tokenizers.modality_config["tok_video_rgb@128"]["tokenizer"]
# A real clip: [0, 1] pixels, resized to 128x128, then normalized to
# roughly [-1, 1] (mean=0.5, std=0.5)
video = torch.randn(1, 3, 17, 128, 128).to(tokenizers.device) # B C T H W, 17 frames
with torch.no_grad():
# Encode: video -> discrete RGB tokens
_, reg_log = rgb_tokenizer.encode(video, return_reg_log=True)
tokens = reg_log["indices"].reshape(-1)
print(tokens.shape) # torch.Size([1280])
# Decode: tokens -> reconstructed video
latents = rgb_tokenizer.regularization.indices_to_codes(reg_log["indices"])
reconstructed = rgb_tokenizer.decoder(latents)
print(reconstructed.shape) # torch.Size([1, 3, 17, 128, 128])
If you would like to tokenize your own multimodal video dataset with these tokenizers, please see README_TOKENIZATION.md. Using these tokenizers, we have also publicly released the tokenized K600-MM dataset, which we used to pretrain Video-4M.
If you find these tokenizers useful, please consider citing:
@article{video4m2026,
title={{Any-to-Any Video Modeling: Modeling the World by Traversing Time and Modalities}},
author={Khattak, Muhammad Uzair and Kim, Won Jun and Abbassi, Reza and Havolli, Albias and Murphy, Michael and Zadeh, Amir and Li, Chuan and Naeem, Muhammad Ferjad and Kar, O\u{g}uzhan Fatih and Bachmann, Roman and Atanov, Andrei and Tombari, Federico and Zamir, Amir},
journal={arXiv preprint},
year={2026},
}
Our tokenizers are adapted from VidTok. We thank the authors for releasing their code.