Downloads · 30 days
0
UseItOrLoseIt/vae-vs-ddpm-celeba
vae-vs-ddpm-celeba is a unconditional image generation model from UseItOrLoseIt. Use it for the unconditional image generation task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
Pretrained checkpoints for a from-scratch PyTorch Variational Autoencoder (VAE) and Denoising Diffusion Probabilistic Model (DDPM), both trained on 64×64 CelebA face crops as a head-to-head comparison of the two gener…
Downloads · 30 days
0
Access
Public
Updated Jul 28, 2026
Repo size
142 MB
Likes
0
Public
Click a slice to open those files.
.pt142 MB · 100%
From the Hugging Face model README
Pretrained checkpoints for a from-scratch PyTorch Variational Autoencoder (VAE) and Denoising Diffusion Probabilistic Model (DDPM), both trained on 64×64 CelebA face crops as a head-to-head comparison of the two generative model families.
Both models are implemented from base torch.nn primitives only — no pretrained VAE, no diffusers library.
📎 Full training code, configs, and technical write-up: AhmedAbdAlkreem/vae-vs-ddpm-celeba
| File | Size | Description |
|---|---|---|
vae_best.pt | 34.7 MB | Best VAE checkpoint (v1.2) |
ddpm_best.pt | 107 MB | Best DDPM checkpoint (v1.2) — U-Net with residual blocks + self-attention (8×8, 16×16) + sinusoidal time embedding |
Both checkpoints correspond to v1.2 in the source repo's version history — the current best result after a documented sampler bug fix (v1.0 → v1.1) and architectural refinements (v1.1 → v1.2). See the CHANGELOG for the full iteration history.
| Model | FID ↓ | Inception Score ↑ | Reconstruction PSNR ↑ | Sampling cost |
|---|---|---|---|---|
| VAE | 97.06 | 2.00 ± 0.06 | 22.47 dB | 1 forward pass |
| DDPM | 26.42 | 2.61 ± 0.17 | n/a | 1000 forward passes (full DDPM) |
DDPM produces sharper, more photorealistic faces at a much higher sampling cost; the VAE is near-instant to sample but characteristically blurry — an architectural trait of pixel-wise reconstruction loss, not a training deficiency. Full quantitative and qualitative comparisons (sample grids, latent-space PCA, interpolations, denoising trajectory) are in the GitHub README.
The model architectures are custom (defined in the source repo under src/models/), so these .pt files are raw state_dict weights, not a transformers/diffusers-compatible model. To load them:
git clone https://github.com/AhmedAbdAlkreem/vae-vs-ddpm-celeba
cd vae-vs-ddpm-celeba
pip install -r requirements.txt
import torch
from huggingface_hub import hf_hub_download
from src.models.vae import VAE
from src.models.unet import UNet # DDPM denoiser
# Download checkpoints from this Hub repo
vae_ckpt = hf_hub_download("UseItOrLoseIt/vae-vs-ddpm-celeba", "vae_best.pt")
ddpm_ckpt = hf_hub_download("UseItOrLoseIt/vae-vs-ddpm-celeba", "ddpm_best.pt")
vae = VAE()
vae.load_state_dict(torch.load(vae_ckpt, map_location="cpu"))
vae.eval()
unet = UNet()
unet.load_state_dict(torch.load(ddpm_ckpt, map_location="cpu"))
unet.eval()
For ready-made sampling and inference scripts (VAE reconstruction, DDPM full/DDIM sampling, latent interpolation), use sample.py and inference.py from the source repo directly — they handle config loading, device placement, and output saving for you.
CelebFaces Attributes (CelebA) Dataset, aligned/cropped to 64×64. v1.2 was trained on a 40,000-image subset (VAE: 40 epochs, DDPM: 50 epochs).
MIT — see LICENSE in the source repo.