Downloads · 30 days
0
HumanAIGC-Team/UCM
UCM is a machine learning model from HumanAIGC-Team. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
0
Access
Public
Updated Aug 7, 2026
Repo size
5.7 GB
Likes
0
Public
Click a slice to open those files.
.ckpt5.7 GB · 100%
From the Hugging Face model README
Tian-Xing Xu^∗¹^, Zi-Xuan Wang^∗¹^, Guangyuan Wang^∗†²^, Li Hu^†²^, Zhongyi Zhang^³^, Peng Zhang^²^, Bang Zhang^‡²^, Song-Hai Zhang^‡¹^
¹Tsinghua University ²Alibaba ³UCST
<sup>∗</sup>Co-first authors <sup>†</sup>Project leaders <sup>‡</sup>Corresponding authors
<br/><img width="320" alt="teaser1" src="assets/teaser.gif"/> <img width="320" alt="teaser2" src="assets/teaser2.gif"/>
</div> <!-- ## 🔆 Notice <!-- We recommend that everyone use English to communicate on issues, as this helps developers from around the world discuss, share experiences, and answer questions together. For further implementation details, please contact `[email protected]` or `[email protected]`. For business licensing and other related inquiries, don't hesitate to contact `[email protected]`. --> <!-- If you find UCM useful, **please help ⭐ this repo**, which is important to Open-Source projects. Thanks! -->We present UCM, a novel framework to explore 4D world of a reference image following the user-specified camera trajectory, which unifies long-term memory and precise camera control via a time-aware positional encoding warping mechanism.
Release Notes:
[2026/08/07] 🔥🔥🔥UCM is released now, have fun!git clone --recursive https://github.com/HumanAIGC/UCM.git
<!-- TODO -->
pip install -r requirements.txt
| Model | Download Links |
|---|---|
| UCM | 🤗 HuggingFace 🤖 ModelScope |
Download models using huggingface-cli:
pip install "huggingface_hub[cli]"
huggingface-cli download HumanAIGC-Team/UCM --local-dir ./workspace/pretrained/
Download models using modelscope-cli:
pip install modelscope
modelscope download --model DAMOXR/UCM --local_dir ./workspace/pretrained/
Run inference code on our provided demo videos, which requires a GPU with ~42GB memory and ~5min to generate a 12s video (241 frames):
python main.py \
--img_path examples/images/frame_0000.png \
--traj_path examples/cameras/cameras_0000.json \
--prompt "The video captures a serene and picturesque scene of a traditional Dutch village on a bright, sunny day. The sky is a vibrant blue with scattered white clouds, creating a perfect backdrop for the charming architecture and lush greenery. The camera pans slowly across the village, revealing a row of quaint houses with red-tiled roofs and brick facades, typical of Dutch design. Some houses have green-painted wooden shutters and doors, adding a touch of color to the scene. A narrow cobblestone street runs through the village, lined with parked cars on both sides, indicating a peaceful residential area."
To obtain all demo videos, you can use the following instruction:
python main.py --metafile examples/examples.csv
Parameters
--img_path: Path to your reference image.--traj_path: Path to your specific camera trajectory file (.json).--prompt: Text prompt.--save_folder: Path to your folder for saving generated videos.--camera_scale_factor: Scales the camera center within the trajectory to match the scale of the 3D scene representation.--num_denoising_steps: The number of denoising iterations. 20 reaches a balance for the Gradio demo. 50 is used in our paper.--guidance_scale: Classifier-Free Guidance scale. The default value 5.0 is recommended.--seed: Seed for initializing the random number generator, controlling the randomness of Gaussian noise sampling.--duration: Only the first Duration frames of the camera trajectory will be processed. -1 represents the whole trajectory.gradio app.py
We have used codes from other great research work, including STream3R and Wan2.1. We sincerely thank the authors for their awesome works!
This project is licensed under the Apache License 2.0.
If you find this work helpful, please consider citing:
@article{xu2026ucm,
title={UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models},
author={Xu, Tianxing and Wang, Zixuan and Wang, Guangyuan and Hu, Li and Zhang, Zhongyi and Zhang, Peng and Zhang, Bang and Zhang, Songhai},
journal={arXiv preprint arXiv:2602.22960},
year={2026}
}
<!-- TODO -->