Downloads · 30 days
0
gpstracquadanio/reconsplat
reconsplat is a image-to-3d model from gpstracquadanio. Use it for the image-to-3d task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as apache-2.0.
Giuseppe Stracquadanio, Kevin Raj, Julia Grabinski, Stefan Roth · ECCV 2026
Downloads · 30 days
0
Access
Public
Updated Sep 5, 2026
Repo size
52.2 GB
Likes
1
Public
Click a slice to open those files.
.ckpt52.2 GB · 100%
From the Hugging Face model README
Giuseppe Stracquadanio, Kevin Raj, Julia Grabinski, Stefan Roth · ECCV 2026
Paper · Project page · Code
ReconSplat reconstructs a scene from a handful of posed images, synthesizing novel views and depth inside and outside the observed frusta. It pairs a feed-forward Gaussian reconstruction with a multi-view latent diffusion model that jointly generates appearance and geometry, so unobserved regions are completed rather than left empty.
This repository holds the released weights. Two stages are trained, and running the model needs one checkpoint from each:
Below we describe our checkpoint hierarchy. The "Init. from" column denotes the weights we use for initializing the corresponding checkpoint. N is the number of context (input) views, M the number of target views, f the VAE downsampling
factor relating image to latent resolution, and B<sub>eff</sub> the effective batch size. All files live under stage2-diffusion/.
| Checkpoint | Init. from | Latent res. | N | M | B<sub>eff</sub> | Steps | Trainable |
|---|---|---|---|---|---|---|---|
| RE10K, 256² | |||||||
re10k-base.ckpt · B<sub>RE10K</sub> | SD 2.1 | 32² | 2 | 4 | 80 | 200K | all |
↳ re10k-upsample-ft.ckpt · + f = 8 → 4 | B<sub>RE10K</sub> | 64² | 2 | 4 | 24 | 50K | all |
↳ re10k-video-ft.ckpt · + Video | B<sub>RE10K</sub> | 32² | 2 | 10 | 32 | 100K | 3D convs. only |
| DL3DV, 256 × 448 | |||||||
dl3dv-base.ckpt · B<sub>DL3DV</sub> | B<sub>RE10K</sub> | 32 × 56 | 4 | 4 | 32 | 140K | all |
↳ dl3dv-upsample-ft.ckpt · + f = 8 → 4 | B<sub>DL3DV</sub> | 64 × 112 | 4 | 4 | 8 | 50K | all |
↳ dl3dv-video-ft.ckpt · + Video | B<sub>DL3DV</sub> | 32 × 56 | 4 | 10 | 24 | 100K | 3D convs. only |
-upsample-ft halves the effective downsampling factor (f = 8 → 4) by bilinearly upsampling
the input images 2×, so the denoiser runs on a 4× larger token grid and recovers finer detail.
Predictions are mapped back to the native resolution.-video-ft fine-tunes on 10-frame sequences for temporally consistent dense camera trajectories. Only the "temporal" residual blocks (i.e., 3D convolutions) are trained; every other parameter stays frozen.[!TIP]
-video-ft: Use these for rendering videos and for point-cloud extraction (recommended).
The table covers only what differs between checkpoints. For the shared training setup, including objective, optimizer and schedule, noise schedule, guidance, and Stage 1 VAE recipe, see the paper.
| File | Size | Pairs with |
|---|---|---|
stage1-vae/re10k.ckpt | 0.7 GB | any re10k-* Stage-2 checkpoint |
stage1-vae/dl3dv.ckpt | 0.7 GB | any dl3dv-* Stage-2 checkpoint |
stage2-diffusion/re10k-base.ckpt | 8.2 GB | stage1-vae/re10k.ckpt |
stage2-diffusion/re10k-upsample-ft.ckpt | 8.2 GB | stage1-vae/re10k.ckpt |
stage2-diffusion/re10k-video-ft.ckpt | 8.9 GB | stage1-vae/re10k.ckpt |
stage2-diffusion/dl3dv-base.ckpt | 8.2 GB | stage1-vae/dl3dv.ckpt |
stage2-diffusion/dl3dv-upsample-ft.ckpt | 8.2 GB | stage1-vae/dl3dv.ckpt |
stage2-diffusion/dl3dv-video-ft.ckpt | 8.9 GB | stage1-vae/dl3dv.ckpt |
poses/re10k/test/ | 50 MB | — (dataset poses, see below) |
poses/dl3dv/test/ | 2.5 MB | — (dataset poses, see below) |
Clone the code and follow its setup instructions (this includes building a custom CUDA rasterization kernel), then download the pair you need:
git clone https://github.com/visinf/reconsplat.git reconsplat
cd reconsplat
pip install -U "huggingface_hub[cli]"
# e.g., to download the base RE10K model.
hf download gpstracquadanio/reconsplat \
stage1-vae/re10k.ckpt stage2-diffusion/re10k-base.ckpt \
--local-dir checkpoints
Each checkpoint has a matching evaluation config, which already points at its default saving location:
python -m src.main +experiment=re10k_diffusion_release_evals \
mode=test \
dataset/view_sampler=evaluation \
dataset.view_sampler.index_path=assets/re10k_evaluation/<index>.json \
test.compute_scores=true
| Checkpoint | +experiment= |
|---|---|
re10k-base | re10k_diffusion_release_evals |
re10k-upsample-ft | re10k_diffusion_release_upsample_ft_evals |
re10k-video-ft | re10k_diffusion_release_video_ft_evals |
dl3dv-base | dl3dv_diffusion_release_ft_evals |
dl3dv-upsample-ft | dl3dv_diffusion_release_upsample_ft_evals |
dl3dv-video-ft | dl3dv_diffusion_release_video_ft_evals |
[!NOTE] To render videos instead of scoring individual views, add
test.save_video=truetest.save_image=falsetest.compute_scores=falseand pointindex_pathat a video evaluation index.
[!NOTE] Stage 2 checkpoints contain both the EMA and the raw denoiser weights. Inference uses the EMA weights by default (
load_ema_weightsalways resolves to true in test mode); you can use the raw weights for further fine-tuning.
To simplify evaluation with VGGT-estimated cameras, we release cached camera predictions for the RE10K and DL3DV test splits. They are stored under poses/re10k/test/ and poses/dl3dv/test/.
Download them with:
mkdir -p datasets # or symlink it to wherever you keep the RE10K/DL3DV data
hf download gpstracquadanio/reconsplat \
--include "poses/re10k/*" \
--local-dir datasets
Then run evaluation with:
dataset.cameras_root=datasets/poses/re10k \
dataset.load_depth_labels=false
For DL3DV, replace re10k with dl3dv.
Apache 2.0. The codebase also incorporates third-party MIT-licensed components, whose copyright notices are retained in the files concerned. Dataset terms remain with their original providers.
@article{stracquadanio2026reconsplat,
title = {ReconSplat: Generalizable 3D scene reconstruction beyond observed views},
author = {Stracquadanio, Giuseppe and Raj, Kevin and Grabinski, Julia and Roth, Stefan},
journal = {{ECCV}},
year = {2026},
}