Downloads · 30 days
0
isno0907/ldmae
ldmae is a machine learning model from isno0907. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Junho Lee\, Jeongwoo Shin\, Hyungwook Choi, Joonseok Lee†
Downloads · 30 days
0
Access
Public
Updated Nov 7, 2025
Repo size
12.9 GB
Likes
2
Public
Click a slice to open those files.
.npz8.7 GB · 67%
From the Hugging Face model README
Junho Lee*, Jeongwoo Shin*, Hyungwook Choi, Joonseok Lee†
Seoul National University, Seoul, Korea * Equal contribution † Corresponding author
This project implements Latent Diffusion Models with Masked AutoEncoders (LDMAE), presented at ICCV 2025. We analyze the role of autoencoders in LDMs and identify three key properties: latent smoothness, perceptual compression quality, and reconstruction quality. We demonstrate that existing autoencoders fail to simultaneously satisfy all three properties, and propose Variational Masked AutoEncoders (VMAEs), taking advantage of the hierarchical features maintained by Masked AutoEncoders. Through comprehensive experiments, we demonstrate significantly enhanced image generation quality and computational efficiency.
The codebase is built upon MAE and LightningDiT.
conda create -n ldmae python=3.10
conda activate ldmae
pip install -r requirements.txt
ldmae_for_github/
├── LDMAE/ # Main diffusion model implementation (based on LightningDiT)
│ ├── configs/ # Configuration files for different datasets
│ ├── datasets/ # Dataset loaders and utilities
│ ├── models/ # Model architectures
│ ├── tokenizer/ # Tokenization modules
│ └── pretrain_weight/ # Directory for pretrained weights
├── VMAE/ # Masked Autoencoder implementation
│ ├── train_ae.sh # Autoencoder training script
│ └── ...
└── requirements.txt # Python dependencies
First, train the autoencoder model using the VMAE module:
cd VMAE
bash train_ae.sh
The training script includes:
After training is complete, save the trained model checkpoint as vmaef8d16.pth in the LDMAE/pretrain_weight/ directory.
** Pretrained checkpoints are also available HERE **
Before proceeding with feature extraction and training, configure the dataset paths in the config files located in the LDMAE/configs/ directory:
configs/imagenet/lightningdit_b_vmae_f8d16_cfg.yamlconfigs/celeba_hq/lightningdit_b_vmae_f8d16_cfg.yamlUpdate the dataset paths according to your local setup.
Extract features from your datasets using the trained autoencoder:
cd LDMAE
bash run_extract_feature.sh configs/imagenet/lightningdit_b_vmae_f8d16_cfg.yaml
cd LDMAE
bash run_extract_feature.sh configs/celeba_hq/lightningdit_b_vmae_f8d16_cfg.yaml
Train the diffusion model on the extracted features:
bash run_train.sh configs/imagenet/lightningdit_b_vmae_f8d16_cfg.yaml
bash run_train.sh configs/celeba_hq/lightningdit_b_vmae_f8d16_cfg.yaml
Generate images using the trained model:
bash run_inference.sh {CONFIG_PATH}
Replace {CONFIG_PATH} with the path to your configuration file (e.g., configs/imagenet/lightningdit_b_vmae_f8d16_cfg.yaml).
python tools/save_npz.py {CONFIG_PATH} # save your npz
python tools/evaluator.py /path/to/reference.npz /path/to/your.npz # calculate metrics
You can download FID stats from HERE
The project includes various configuration files for different model variants and datasets:
LDMAE/configs/imagenet/LDMAE/configs/celeba_hq/Each configuration file specifies:
If you use this code in your research, please cite our paper:
@InProceedings{Lee_2025_ICCV,
author = {Lee, Junho and Shin, Jeongwoo and Choi, Hyungwook and Lee, Joonseok},
title = {Latent Diffusion Models with Masked AutoEncoders},
booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
month = {October},
year = {2025},
pages = {17422-17431}
}
@article{lee2025latent,
title={Latent Diffusion Models with Masked AutoEncoders},
author={Lee, Junho and Shin, Jeongwoo and Choi, Hyungwook and Lee, Joonseok},
journal={arXiv preprint arXiv:2507.09984},
year={2025}
}