Skip to content

nyu-visionx

siglip2_decoder

nyu-visionx/siglip2_decoder

siglip2_decoder is a image-to-image model from nyu-visionx. Use it when you need one image transformed into another. It is set up for transformers. The card lists the license as mit.

This repository contains artifacts related to the paper Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders.

Downloads · 30 days

1.4K

2% of all-time downloads

All-time downloads

73.7K

Public

Repo size

1.7 GB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.pt1.7 GB · 100%

At a glance

Task
Image-to-Image
Library
transformers
License
mit
Model type
vit_mae
Access
Public
Created
Jan 8, 2026
Updated
Jan 24, 2026
SHA
4a728df7
Task
Image-to-Image
Library
transformers
Type
vit_mae
License
mit
Created
Jan 8, 2026
Updated
Jan 24, 2026
siglip2_decoder — AI Model — AIMarketly