Downloads · 30 days
0
yqi19/VIPRA-reproduce
VIPRA-reproduce is a machine learning model from yqi19. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
From-scratch reproduction of VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation (arXiv 2607.23472v1, tech report, no official code release).
Downloads · 30 days
0
Access
Public
Updated Aug 21, 2026
Repo size
1.8 GB
Likes
0
Public
Click a slice to open those files.
.mp41.2 GB · 66%
From the Hugging Face model README
From-scratch reproduction of VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation (arXiv 2607.23472v1, tech report, no official code release).
📖 Main documentation is in Chinese: README_zh.md — it explains
what the paper does, how the method works, and what every pipeline step means.
📊 Results: RESULTS_zh.md — dataset statistics, throughput,
and the engineering issues that had to be fixed to make it run.
VIPER conditions an image-to-video generator on a reference video that demonstrates a physical process, so the physics transfers to a new target scene. A frozen MLLM reads the reference alongside learnable query tokens; their hidden states become physics condition tokens that are concatenated with the text context of a Wan2.2 DiT. Trained in three hierarchical stages.
viper/ model, data pipeline, training, inference, evaluation
scripts/ slurm / torchrun launchers
third_party/ official Wan2.2 code (DiT / VAE / umT5 definitions)
source env.sh # all caches point at lustre; nothing is written to $HOME
sbatch scripts/full_run.sbatch
| File | Role |
|---|---|
viper/physics_encoder.py | learnable queries + frozen Qwen3-VL + 3-layer connector |
viper/wan_viper.py | physics-token injection into the Wan2.2 DiT + LoRA |
viper/data/filter_clips.py | clip filtering (compression / watermark / cuts / motion) |
viper/data/annotate.py | 3-axis physics annotation (material / trajectory / impact) |
viper/data/build_pairs.py | bucketing + MLLM transferability filtering |
viper/train.py | 3-stage hierarchical training |
viper/infer.py, infer_batch.py | Flow-Euler sampling, 50 steps, CFG 6.0 |
viper/eval.py | VBench-style metrics + VLM-as-Judge physical similarity |
viper/eval_loss.py | paired ablation: reference tokens vs zeroed tokens |
results/comparison/<ref_id>__<tgt_id>/Nine held-out validation cases. Each folder contains the full picture:
| File | What it is |
|---|---|
contact_sheet.png | Easiest to read — 4 rows (reference / ground truth / baseline / VIPER) × 6 time samples in one image |
grid.mp4 | The same four videos tiled 2×2 with labels |
reference.mp4 | The reference video — the physical process being transferred |
target_gt.mp4 | The ground-truth target — what a correct answer looks like |
target_img.png | The target image actually fed to the generator (frame 0 of the target) |
baseline.mp4 | Wan2.2 output with NO reference (physics tokens zeroed = untouched base model) |
viper.mp4 | VIPER output — same image + prompt, plus physics tokens from the reference |
info.json | prompt, physics labels, judge score, sampling params, seed |
baseline.mp4 and viper.mp4 use identical seeds, so any difference between
them is attributable to the reference stream alone.
| Path | What it is |
|---|---|
results/eval_comparison.json | VBench-style metrics + VLM-judge physical similarity, per variant (baseline vs viper) over the 9 cases |
results/eval_loss_stage1.json | Paired conditioning ablation: flow-matching loss with reference tokens vs zeroed, same sample/noise/timestep (36 measurements) |
results/comparison_manifest.jsonl | Index feeding viper.eval |
Built from the full csusupergear/WISA-80K-wan480p-16fps-81f subset (276 shards,
27543 clips) via scripts/fetch_wisa.py + scripts/scaleup_{1,2,3}_*.sbatch.
The pilot column is the earlier 20-shard run, kept for reference.
| Path | Pilot | Full | Stage of the pipeline |
|---|---|---|---|
data/videos/ | 2000 | 27543 | downloaded WISA wan480p clips |
data/filtered_clips.jsonl | 523 | 7711 | after clip filtering |
data/annotated.jsonl | 523 | 7693 | after 3-axis physics annotation |
data/pairs_candidates.jsonl | 996 | 15608 | after label bucketing |
data/pairs.jsonl | 64 | 934 | after Qwen3-VL-32B transferability filtering |
data/pairs_train.jsonl / pairs_val.jsonl | 23 / 9 | 553 / 74 | video-disjoint split |
data/clips_train.jsonl | — | 7568 | stage-1 clips minus held-out videos |
What the filters cut, at full scale: 14762 near-static (motion_score < 1.0),
4126 non-mechanical labels, 725 low quality, 219 shot transitions, 0 watermarked.
The 32B judge kept 934 of ~9600 pairs it got through (9.7%).
Only ~9600 of the 15608 candidates were judged: the judge runs at ~6 s/pair/rank
and the 4h interactive limit cuts it off there. --deadline_min makes that stop
graceful so the split still runs; raise it on a longer partition to judge the rest.
logs/ — judge.log (pair filtering), train_stage1_interactive.log (training),
comparison.log (video generation), eval_comparison.log, eval_loss.log.
Base model weights (models/, ~115GB — fetch from Wan-AI/Wan2.2-I2V-A14B,
Qwen/Qwen3-VL-4B-Instruct, Qwen/Qwen3-VL-32B-Instruct), our trained
checkpoints, the source video corpus (data/videos/, from
csusupergear/WISA-80K-wan480p-16fps-81f), and the precomputed latent cache.
⚠️ Read
RESULTS_zh.mdbefore interpreting the videos. Only 150 stage-1 steps were trained (~1% of the paper's 15K budget), soviper.mp4andbaseline.mp4are near-identical by design — the zero-init connector has not yet grown enough to influence the DiT. These are an infrastructure baseline, not a validation of the method.