Downloads · 30 days
0
zhaoyian01/MiniWorld
MiniWorld is a robotics model from zhaoyian01. Use it for the robotics task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as apache-2.0.
MiniWorld: Democratizing the Training of Video World Models from Scratch
Downloads · 30 days
0
Access
Public
Updated Aug 4, 2026
Repo size
11.9 GB
Likes
3
Trending 1
Click a slice to open those files.
.pt11.9 GB · 100%
From the Hugging Face model README
MiniWorld: Democratizing the Training of Video World Models from Scratch
<div align="center" style="line-height: 1;"> <a href="https://zhao-yian.github.io/MiniWorld/" target="_blank" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img alt="Project Page" src="https://img.shields.io/badge/Project-Page-1f6feb?style=for-the-badge&logo=googlechrome&logoColor=white" style="display: block;"/> </a> <a href="https://arxiv.org/abs/2608.01127" target="_blank" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img alt="arXiv" src="https://img.shields.io/badge/arXiv-2608.01127-b31b1b?style=for-the-badge&logo=arxiv&logoColor=white" style="display: block;"/> </a> <a href="https://github.com/zhao-yian/MiniWorld" target="_blank" style="display: inline-block; vertical-align: middle; margin: 2px;"> <img alt="GitHub" src="https://img.shields.io/badge/GitHub-Code-181717?style=for-the-badge&logo=github&logoColor=white" style="display: block;"/> </a> </div>MiniWorld is a minimal and reproducible framework for training streaming video world models from scratch. Instead of adapting a pretrained bidirectional video generator, it directly learns causal next-state prediction with a block-causal Video Diffusion Transformer and Rectified Flow.
The same architecture supports two control modalities:
This Hugging Face repository hosts the MiniWorld model checkpoints. Code, training scripts, and evaluation utilities live in the GitHub repository.
Each tile is a 253-frame streaming rollout from the 1B checkpoint, generated
from a single observed frame plus the control signal.
MiniWorld uses a block-causal Video Diffusion Transformer trained with Rectified Flow in the latent space of the Wan2.2 VAE. During inference, MiniWorld performs streaming generation with a rolling KV cache and pipelined asynchronous denoising, enabling long-horizon generation under bounded online computation.
Key components:
The complete model can be trained in several days on a single 8-GPU server.
Sampling requires matching the checkpoint with the corresponding dataset and model scale.
| Dataset | Model | Status | Checkpoint |
|---|---|---|---|
| DROID | MiniWorld-0.5B | Available | MiniWorld_0_5b_droid.pt |
| DROID | MiniWorld-1B | Available | MiniWorld_1b_droid.pt |
| DROID | MiniWorld-3B | Coming soon | -- |
| RealEstate10K | MiniWorld-0.5B | Available | MiniWorld_0_5b_re10k.pt |
| RealEstate10K | MiniWorld-1B | Available | MiniWorld_1b_re10k.pt |
| RealEstate10K | MiniWorld-3B | Coming soon | -- |
Download a single checkpoint with:
hf download zhaoyian01/MiniWorld \
--include "MiniWorld_1b_droid.pt" \
--local-dir checkpoints/miniworld
MODEL is the identifier expected by the training and sampling scripts in the
GitHub repository.
| Model | MODEL | Depth | Width | Heads | Parameters |
|---|---|---|---|---|---|
| MiniWorld-B | B | 12 | 768 | 12 | 0.12B |
| MiniWorld-L | L | 24 | 1024 | 16 | 0.39B |
| MiniWorld-0.5B | 0.5B | 28 | 1152 | 16 | 0.55B |
| MiniWorld-1B | 1B | 28 | 1536 | 12 | 1B |
| MiniWorld-3B | 3B | 32 | 2560 | 20 | 3B |
MiniWorld is intended for research on streaming video world models, including:
MiniWorld is a research baseline and is not intended as a general-purpose text-to-video model.
Inference requires the MiniWorld codebase and the pretrained Wan2.2 VAE:
Wan-AI/Wan2.2-TI2V-5BDownload the VAE:
hf download Wan-AI/Wan2.2-TI2V-5B \
--include "Wan2.2_VAE.pth" \
--local-dir checkpoints/wan2.2
Clone the MiniWorld codebase, install its requirements, then download the desired checkpoint. All commands are run from the repository root.
The default sampler uses one observed frame as initial context, eight in-flight
chunks and a 24-chunk rolling KV cache (a 64-frame active attention window), one
persistent sink frame, 100 denoising steps with classifier-free guidance at
scale 2.0, and a 64-latent-frame rollout corresponding to 253 RGB frames.
Generated videos are saved to ${SAMPLE_DIR}/pred/.
DATA_ROOT=/path/to/droid_lerobot \
CKPT=/path/to/MiniWorld_1b_droid.pt \
VAE_CKPT=checkpoints/wan2.2/Wan2.2_VAE.pth \
MODEL=1B \
bash scripts/sample_droid.sh
DATA_ROOT=/path/to/re10k/videos \
POSE_DIR=/path/to/re10k/poses \
CKPT=/path/to/MiniWorld_1b_re10k.pt \
VAE_CKPT=checkpoints/wan2.2/Wan2.2_VAE.pth \
MODEL=1B \
bash scripts/sample_re10k.sh
GPU=0 \
TOTAL_LEN=96 \
CFG_SCALE=2.0 \
SAMPLE_NUM_VIDEOS=10 \
STREAM_INFLIGHT_CHUNKS=8 \
STREAM_MAX_CACHE_CHUNKS=24 \
STREAM_SINK_SIZE=1 \
bash scripts/sample_droid.sh
TOTAL_LEN sets the rollout length in latent frames and can exceed the trained
window, since streaming keeps the attention span bounded; TOTAL_LEN=96 yields
381 RGB frames from a 64-frame checkpoint. MiniWorld is a streaming model and
does not assume a fixed generation horizon.
A RealEstate10K checkpoint can also animate a single image along a procedural camera trajectory, without any dataset on disk:
PYTHONPATH=. python -m miniworld.sample \
--dataset re10k \
--init_image /path/to/first_frame.png \
--custom_camera_trajectory orbit_right \
--checkpoint /path/to/MiniWorld_1b_re10k.pt \
--vae_checkpoint checkpoints/wan2.2/Wan2.2_VAE.pth \
--sample_dir samples/re10k_orbit_right \
--wm_model 1B \
--total_len 64 \
--sample_num_videos 1 \
--trajectory_magnitude 3.0
These checkpoints are trained on raw (unnormalized) translations, so
--trajectory_magnitude is worth tuning: 1.0 is almost static, 3.0 is a
good default at --total_len 64, and values above 5.0 degrade the second half
of the rollout. Scale it with the rollout length to keep the same apparent
speed. See the GitHub README for the full list of trajectories.
MiniWorld is a research model trained and evaluated at modest resolution and on limited domains. It may exhibit long-horizon drift, geometric errors, temporal inconsistencies, and failures under out-of-distribution actions, poses, scenes, or camera motions. It should not be used for safety-critical simulation or as a faithful physical simulator.
The paper is available on arXiv: arXiv:2608.01127.
If you find MiniWorld useful in your research, please cite:
@article{zhao2026miniworld,
title = {MiniWorld: Democratizing the Training of Video World Models from Scratch},
author = {Zhao, Yian and Zheng, Ruochong and Guo, Hongcan and Yan, Yu and Zhang, Jian and Chen, Jie},
journal = {arXiv preprint arXiv:2608.01127},
year = {2026}
}
These checkpoints are released under the Apache 2.0 license. Please also follow the licenses and usage terms of the underlying datasets (DROID, RealEstate10K) and of the Wan2.2 VAE.