Downloads · 30 days
0
qihoo360/RPiAE
RPiAE is a machine learning model from qihoo360. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<p align="center" <img src="assets/logo.png" alt="RPiAE Logo" width="200" </p
Downloads · 30 days
0
Access
Public
Updated Apr 9, 2026
Repo size
23.9 GB
Likes
4
Public
Click a slice to open those files.
.pt23.8 GB · 100%
From the Hugging Face model README
This repository contains the PyTorch implementation of RPiAE and the corresponding latent diffusion training pipeline.
RPiAE follows a two-stage pipeline:
uv:
conda create -n rpiae python=3.10 -y
conda activate rpiae
pip install uv
# Install PyTorch 2.8.0 with CUDA 12.9 # or your own cuda version
uv pip install torch==2.8.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu129
# Install other dependencies
uv pip install -r requirements.txt
Pretrained checkpoints are available on Hugging Face.
Download the pretrain weights to ./model_weights:
hf download qihoo360/RPiAE \
--repo-type model \
--local-dir ./model_weights
--data-path.All training and sampling entrypoints are driven by OmegaConf YAML files. A single config describes the Stage 1 autoencoder, the Stage 2 diffusion model, and the solver used during training or inference. A minimal example looks like:
stage_1:
target: stage1.RPiAE / stage1.RPiAE_VB
params: { ... }
ckpt: <path_to_ckpt>
stage_2:
target: stage2.models.lightningDiT.LightningDiT
params: { ... }
ckpt: <path_to_ckpt>
transport:
params:
path_type: Linear
prediction: velocity
...
sampler:
mode: ODE
params:
num_steps: 50
...
guidance:
method: cfg/autoguidance
scale: 1.0
...
misc:
latent_size: [64, 16, 16]
num_classes: 1000
training:
...
eval:
...
stage_1 defines the RPiAE training process (reconstruction-oriented training).stage_2 defines the generation model (LightningDiT) in the RPiAE latent space.transport, sampler, and guidance control ODE/SDE solving and guidance strategy.misc stores latent shape and shared constants.training and eval contain optimization and online evaluation settings.gan block for discriminator and loss schedule.configs/stage1/pretrained/DINOv2-B_decXL_RPiAE.yamlconfigs/stage1/training/DINOv2-B_decXL_RPiAE_stage1.yamlconfigs/stage1/training/DINOv2-B_decXL_RPiAE_stage2.yamlconfigs/stage1/training/DINOv2-B_decXL_RPiAE_stage3.yamlconfigs/stage2/training/ImageNet256/LightingDiT-XL_f16d64rpiae-v2b_vitxl.yamlconfigs/stage2/sampling/ImageNet256/LightingDiT-XL_d64rpiae-v2b_vitxl.yamlconfigs/stage2/sampling/ImageNet256/LightingDiT-XL_d64rpiae-v2b_vitxl_AG.yamlUse the provided shell scripts with the corresponding configs.
If you use wandb logging, also configure:
EXPERIMENT_NAME=
ENTITY=
PROJECT=
bash run_train_stage1_rpiae_s1.sh \
configs/stage1/training/DINOv2-B_decXL_RPiAE_stage1.yaml
bash run_train_stage1_rpiae_s23.sh \
configs/stage1/training/DINOv2-B_decXL_RPiAE_stage2.yaml
bash run_train_stage1_rpiae_s23.sh \
configs/stage1/training/DINOv2-B_decXL_RPiAE_stage3.yaml
For multi-GPU or multi-node training, please use the *_mult*.sh scripts instead.
Before launching, configure the distributed variables in the shell script:
RANK=
MASTER_ADDR=
GPUS_PER_NODE=
NNODES=
MASTER_PORT=
Then run the corresponding multi-GPU scripts.
bash run_train_stage1_mult_rpiae_s1.sh \
configs/stage1/training/DINOv2-B_decXL_RPiAE_stage1.yaml
bash run_train_stage1_mult_rpiae_s23.sh \
configs/stage1/training/DINOv2-B_decXL_RPiAE_stage2.yaml
bash run_train_stage1_mult_rpiae_s23.sh \
configs/stage1/training/DINOv2-B_decXL_RPiAE_stage3.yaml
bash run_sample_reconstruction_eval.sh \
configs/stage1/pretrained/DINOv2-B_decXL_RPiAE.yaml
torchrun --standalone --nnodes=1 --nproc_per_node=N \
src/train_diffusion_rpiae.py \
--config <training_config> \
--data-path <imagenet_train_split> \
--results-dir ckpts/diffusion \
--compile \
--precision fp32
For multi-GPU / multi-node training, use:
bash run_train_mult_diffusion.sh \
configs/stage2/training/ImageNet256/LightingDiT-XL_f16d64rpiae-v2b_vitxl.yaml
For multi-node launch, set:
export RANK=<node_rank>
export MASTER_ADDR=<master_node_ip_or_hostname>
Although bf16 is supported, we recommend using fp32 for more stable training.
src/sample.py uses the same config schema to draw a small batch of images on a
single device and saves them to sample.png:
python src/sample.py \
--config <sample_config> \
--seed 42
src/sample_ddp.py parallelises sampling across GPUs, producing PNGs and an
FID-ready .npz:
torchrun --standalone --nnodes=1 --nproc_per_node=N \
src/sample_ddp.py \
--config <sample_config> \
--sample-dir samples \
--precision fp32/bf16 \
--label-sampling equal
--label-sampling {equal,random}: equal uses exactly 50 images per class for FID-50k; random uniformly samples labels. We use equal by default. We recommend using fp32 when model FID is low.
Autoguidance and classifier-free guidance are controlled via the config’s guidance block.
Use the ADM evaluation suite to score generated samples:
Clone the repo:
git clone https://github.com/openai/guided-diffusion.git
cd guided-diffusion/evaluation
Create an environment and install dependencies:
conda create -n adm-fid python=3.10
conda activate adm-fid
pip install 'tensorflow[and-cuda]'==2.19 scipy requests tqdm
Download ImageNet statistics (256×256 shown here):
wget https://openaipublic.blob.core.windows.net/diffusion/jul-2021/ref_batches/imagenet/256/VIRTUAL_imagenet256_labeled.npz
Evaluate:
python evaluator.py VIRTUAL_imagenet256_labeled.npz /path/to/samples.npz
This code is built upon the following repositories:
If you find this repository useful, please consider citing our paper:
@misc{RPiAE,
title={RPiAE: A Representation-Pivoted Autoencoder Enhancing Both Image Generation and Editing},
author={Yue Gong and Hongyu Li and Shanyuan Liu and Bo Cheng and Yuhang Ma and Liebucha Wu and Xiaoyu Wu and Manyuan Zhang and Dawei Leng and Yuhui Yin and Lijun Zhang},
year={2026},
eprint={2603.19206},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2603.19206},
}