Downloads · 30 days
1.4K
2% of all-time downloads
nyu-visionx/siglip2_decoder
siglip2_decoder is a image-to-image model from nyu-visionx. Use it when you need one image transformed into another. It is set up for transformers. The card lists the license as mit.
This repository contains artifacts related to the paper Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders.
Downloads · 30 days
1.4K
2% of all-time downloads
All-time downloads
73.7K
Public
Repo size
1.7 GB
Likes
0
Public
Click a slice to open those files.
.pt1.7 GB · 100%
From the Hugging Face model README
This repository contains artifacts related to the paper Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders.
Representation Autoencoders (RAEs) provide a simplified and powerful alternative to VAEs for large-scale text-to-image generation. Scale-RAE demonstrates that training diffusion models in high-dimensional semantic latent spaces (using encoders like SigLIP-2) leads to faster convergence, better generation quality, and improved stability compared to state-of-the-art VAE-based foundations.
For detailed instructions on installation, training, and inference, please visit the official GitHub repository.
This decoder is also directly compatitable with original RAE codebase. Try it out by simply swapping the encoder with google/siglip2-so400m-patch14-224!
@article{scale-rae-2026,
title={Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders},
author={Shengbang Tong and Boyang Zheng and Ziteng Wang and Bingda Tang and Nanye Ma and Ellis Brown and Jihan Yang and Rob Fergus and Yann LeCun and Saining Xie},
journal={arXiv preprint arXiv:2601.16208},
year={2026}
}