Downloads · 30 days
0
Rishu011/Matrix-Game-3.0
Matrix-Game-3.0 is a image-text-to-video model from Rishu011. Use it for the image-text-to-video task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as apache-2.0.
<div style="display: flex; justify-content: center; gap: 10px;" <a href="https://github.com/SkyworkAI/Matrix-Game" <img src="https://img.shields.io/badge/GitHub-100000?style=flat&logo=github&logoColor=white" alt="GitH…
Downloads · 30 days
0
Access
Public
Updated Mar 30, 2026
Repo size
56.6 GB
Likes
1
Public
Click a slice to open those files.
.safetensors38.8 GB · 69%
From the Hugging Face model README
Matrix-Game-3.0 is an open-sourced, memory-augmented interactive world model designed for 720p real-time long-form video generation.
Our framework unifies three stages into an end-to-end pipeline:

Create a conda environment and install dependencies:
conda create -n matrix-game-3.0 python=3.12 -y
conda activate matrix-game-3.0
# install FlashAttention
# Our project also depends on [FlashAttention](https://github.com/Dao-AILab/flash-attention)
git clone https://github.com/SkyworkAI/Matrix-Game-3.0.git
cd Matrix-Game-3.0
pip install -r requirements.txt
pip install "huggingface_hub[cli]"
huggingface-cli download Matrix-Game-3.0 --local-dir Matrix-Game-3.0
Before running inference, you need to prepare:
After downloading pretrained models, you can use the following command to generate an interactive video with random actions:
torchrun --nproc_per_node=$NUM_GPUS generate.py --size 704*1280 --dit_fsdp --t5_fsdp --ckpt_dir Matrix-Game-3.0 --fa_version 3 --use_int8 --num_iterations 12 --num_inference_steps 3 --image demo_images/000/image.png --prompt "a vintage gas station with a classic car parked under a canopy, set against a desert landscape." --save_name test --seed 42 --compile_vae --lightvae_pruning_rate 0.5 --vae_type mg_lightvae --output_dir ./output
# "num_iterations" refers to the number of iterations you want to generate. The total number of frames generated is given by:57 + (num_iterations - 1) * 40
Tips:
If you want to use the base model, you can use "--use_base_model --num_inference_steps 50". Otherwise if you want to generating the interactive videos with your own input actions, you can use "--interactive".
With multiple GPUs, you can pass --use_async_vae --async_vae_warmup_iters 1 to speed up inference.
If you find this work useful for your research, please kindly cite our paper:
@misc{2026matrix,
title={Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory},
author={{Skywork AI Matrix-Game Team}},
year={2026},
howpublished={Technical report},
url={https://github.com/SkyworkAI/Matrix-Game/blob/main/Matrix-Game-3/assets/pdf/report.pdf}
}