Downloads · 30 days
0
zhen-nan/DiP
DiP is a unconditional image generation model from zhen-nan. Use it for the unconditional image generation task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<div style="text-align: center;" <a href="https://arxiv.org/abs/2511.18822"<img src="https://img.shields.io/badge/arXiv-2511.18822-b31b1b.svg" alt="arXiv"</a </div
Downloads · 30 days
0
Access
Public
Updated Jun 10, 2026
Repo size
10.2 GB
Likes
15
Public
Click a slice to open those files.
.ckpt10.2 GB · 100%
From the Hugging Face model README
Diffusion models face a fundamental trade-off between generation quality and computational efficiency. Latent Diffusion Models (LDMs) offer an efficient solution but suffer from potential information loss and non-end-to-end training. In contrast, existing pixel space models bypass VAEs but are computationally prohibitive for high-resolution synthesis. To resolve this dilemma, we propose DiP, an efficient pixel space diffusion framework. DiP decouples generation into a global and a local stage: a Diffusion Transformer (DiT) backbone operates on large patches for efficient global structure construction, while a co-trained lightweight Patch Detailer Head leverages contextual features to restore fine-grained local details. This synergistic design achieves computational efficiency comparable to LDMs without relying on a VAE. DiP is accomplished with up to 10x faster inference speeds than previous method while increasing the total number of parameters by only 0.3%, and achieves an 1.79 FID score on ImageNet 256x256.