Downloads · 30 days
0
QuintonSung/WorldWander
WorldWander is a machine learning model from QuintonSung. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<div align="center" <h1 WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation </h1 </div
Downloads · 30 days
0
Access
Public
Updated Jul 13, 2026
Repo size
3.2 GB
Likes
0
Public
Click a slice to open those files.
.ckpt3.2 GB · 100%
From the Hugging Face model README
Video diffusion models have recently achieved remarkable progress in realism and controllability. However, achieving seamless video translation across different perspectives, such as first-person (egocentric) and third-person (exocentric), remains underexplored. Bridging these perspectives is crucial for filmmaking, embodied AI, and world models.
Motivated by this, we present <b>WorldWander</b>, an in-context learning framework tailored for translating between egocentric and exocentric worlds in video generation. Building upon advanced video diffusion transformers, WorldWander integrates (i) <i>In-Context Perspective Alignment</i> and (ii) <i>Collaborative Position Encoding</i> to efficiently model cross-view synchronization.
Overall framework is shown below:

👋 If you find this code useful for your research, we would appreciate it if you could cite:
@article{song2025worldwander,
title={WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation},
author={Song, Quanjian and Song, Yiren and Peng, Kelly and Gao, Yuan and Shou, Mike Zheng},
journal={arXiv preprint arXiv:2511.22098},
year={2025}
}