Downloads · 30 days
0
nvidia/PixelDiT-ImageNet
PixelDiT-ImageNet is a unconditional image generation model from nvidia. Use it for the unconditional image generation task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as other.
<p align="center" <img src="https://raw.githubusercontent.com/NVlabs/PixelDiT/master/assets/pixeldit-logo.png" height="60" / </p
Downloads · 30 days
0
Access
Public
Updated Apr 15, 2026
Repo size
51.2 GB
Likes
15
Public
Click a slice to open those files.
.ckpt51.2 GB · 100%
From the Hugging Face model README
| Checkpoint | Resolution | Epochs | gFID | CFG Scale | Time Shift | CFG Interval |
|---|---|---|---|---|---|---|
imagenet256_pixeldit_xl_epoch80.ckpt | 256x256 | 80 | 2.36 | 3.25 | 1.0 | [0.1, 1.0] |
imagenet256_pixeldit_xl_epoch160.ckpt | 256x256 | 160 | 1.97 | 3.25 | 1.0 | [0.1, 1.0] |
imagenet256_pixeldit_xl_epoch320.ckpt | 256x256 | 320 | 1.61 | 2.75 | 1.0 | [0.1, 0.9] |
imagenet512_pixeldit_xl.ckpt | 512x512 | 850 | 1.81 | 3.5 | 2.0 | [0.1, 1.0] |
All evaluations use FlowDPMSolver with 100 steps. 50K samples. Metrics follow the ADM evaluation protocol.
pip install -r requirements.txt
cd c2i/
# ImageNet 256x256 (epoch 320, best FID)
torchrun --nproc_per_node=8 main.py predict \
-c configs/pix256_xl.yaml \
--ckpt_path=imagenet256_pixeldit_xl_epoch320.ckpt \
--model.diffusion_sampler.class_path=src.diffusion.FlowDPMSolverSampler \
--model.diffusion_sampler.init_args.num_steps=100 \
--model.diffusion_sampler.init_args.guidance=2.75 \
--model.diffusion_sampler.init_args.timeshift=1.0 \
--model.diffusion_sampler.init_args.guidance_interval_min=0.1 \
--model.diffusion_sampler.init_args.guidance_interval_max=0.9 \
--per_run_seed=false --seed_everything=1000
# ImageNet 512x512
torchrun --nproc_per_node=8 main.py predict \
-c configs/pix512_xl.yaml \
--ckpt_path=imagenet512_pixeldit_xl.ckpt \
--model.diffusion_sampler.class_path=src.diffusion.FlowDPMSolverSampler \
--model.diffusion_sampler.init_args.num_steps=100 \
--model.diffusion_sampler.init_args.guidance=3.5 \
--model.diffusion_sampler.init_args.timeshift=2.0 \
--model.diffusion_sampler.init_args.guidance_interval_min=0.1 \
--model.diffusion_sampler.init_args.guidance_interval_max=1.0 \
--per_run_seed=false --seed_everything=10000
After generating samples, compute FID with the ADM evaluation toolkit.
@inproceedings{yu2025pixeldit,
title={PixelDiT: Pixel Diffusion Transformers for Image Generation},
author={Yongsheng Yu and Wei Xiong and Weili Nie and Yichen Sheng and Shiqiu Liu and Jiebo Luo},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year={2026},
}
This model is released under the NSCLv1 License. The work and any derivative works may only be used for non-commercial (research or evaluation) purposes.