Downloads · 30 days
0
luuuulinnnn/evs-scene-extrapolation
evs-scene-extrapolation is a machine learning model from luuuulinnnn. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A generative video diffusion model for posed scene extrapolation: given a short observed video and a target camera trajectory, generate the video along that trajectory — extrapolating into unobserved regions while sta…
Downloads · 30 days
0
Access
Public
Updated Jul 14, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md3.1 KB · 67%
From the Hugging Face model README
A generative video diffusion model for posed scene extrapolation: given a short observed video and a target camera trajectory, generate the video along that trajectory — extrapolating into unobserved regions while staying consistent on revisit (when the camera returns to previously-seen content).
Built on Wan2.1-VACE-14B, trained with Diffusion Forcing, and conditioned via Plücker camera control + a geometry-aware "sensor token" design.
Diffusion Forcing (in-context memory). Independent per-frame noise levels; the observed frames sit in the same latent sequence at noise ≈ 0 (clean) while target frames are denoised. Memory is in-context — the model's self-attention reads the clean observed frames; no separate memory module.
Sensor token (the contribution). Each token carries, beyond its content latent:
Camera control. PermaVid-style Plücker SimpleAdapter, additive at the patch embedding (not via the
VACE control branch). Per-frame timesteps require no architecture surgery (Wan DiTBlock already accepts
per-token modulation). All new modules are zero-init, so an untrained model equals the base Wan T2V model.
RealEstate10K (static real-estate walkthroughs, smooth cinematic camera). Clips are built as smooth out-and-back trajectories over consecutive frames (forward then retrace) for a genuine revisit with real GT, plus a forward-split eval mode (observe a prefix, extrapolate the forward continuation into unseen). Geometry (ray codes + evidence) is a pure function of the poses and is precomputed/cached.
code/
evs_model.py sensor-token model (ray-PE + evidence + camera + Diffusion Forcing)
cache_re10k.py RealEstate10K -> cached clips (latents + poses + ray + evidence)
cache_smooth.py DL3DV smooth out-and-back builder (earlier data source)
train_df.py FSDP trainer (per-frame DF noise, target-masked flow-matching, resume + safe save)
infer_df.py Diffusion-Forcing sampling (clean context + generated target)
make_*_viz.py qualitative viz (revisit / data / forward-split comparisons)
make_eval_metrics.py forward-split extrapolation metrics (PSNR/SSIM on generated frames)
*.sbatch, evs_fsdp*.yaml SLURM + accelerate-FSDP launch configs
docs/
DESIGN.md, BUILD_LOG.md design rationale + running build log
Work in progress. Model + pipeline implemented and verified; training on RealEstate10K. Checkpoints and qualitative samples: see the companion HuggingFace repo.