Downloads ยท 30 days
2
2% of all-time downloads
leonardPKU/DnD-Transformer
DnD-Transformer is a machine learning model from leonardPKU. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<p align="center" ๐ค <a href="https://huggingface.co/leonardPKU/DnD-Transformer"Model</a   | ๐ค <a href="" Dataset (Coming Soon)</a  |   ๐ <a href="https://arxiv.org/abs/2410.01912"Paper</a |โฆ
Downloads ยท 30 days
2
2% of all-time downloads
All-time downloads
98
Public
Repo size
96.1 GB
Likes
3
Public
Click a slice to open those files.
.pt32.8 GB ยท 100%
From the Hugging Face model README
What's New?
A better AR image genenation paradigm and transformer model structure based on 2D autoregression. It generates images of higher quality without increasing computation budget.
A spark of vision-language intelligence for the first time, enabling unconditional rich-text image generation, outperforming diffusion models like DDPM and Stable Diffusion on dedicated rich-text image datasets, highlighting the distinct advantage of autoregressive models for multimodal modeling.
Text-Image
| Code Size | Link |
|---|---|
| 24x24x1 | ๐ค |
ImageNet
| Code Size | Link | rFID |
|---|---|---|
| 16x16x2 | ๐ค | 0.92 |
arXiv-Image
coming soon~
Text-Image
| Code Shape | Model Size | Link |
|---|---|---|
| 24x24x1 | XXL | ๐ค |
ImageNet
arXiv-Image
coming soon~
conda create -n DnD python=3.10
conda activate DnD
pip install -r requirements.txt
Sampling Text-Image Examples
cd ./src
bash ./scripts/sampling_dnd_transformer_text_image.sh # edit the address for vq model checkpoint and dnd-transformer checkpoint
Sampling ImageNet Examples
cd ./src
bash ./scripts/sampling_dnd_transformer_imagenet.sh # edit the address for vq model checkpoint and dnd-transformer checkpoint
# An npz would be saved after genearting 50k images, you can follow https://github.com/openai/guided-diffusion/tree/main/evaluations to compute the generated FID.
Training code and Dataset are coming soon!
@misc{chen2024sparkvisionlanguageintelligence2dimensional,
title={A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation},
author={Liang Chen and Sinan Tan and Zefan Cai and Weichu Xie and Haozhe Zhao and Yichi Zhang and Junyang Lin and Jinze Bai and Tianyu Liu and Baobao Chang},
year={2024},
eprint={2410.01912},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2410.01912},
}