Downloads · 30 days
0
Zyriix/prologue
prologue is a unconditional image generation model from Zyriix. Use it for the unconditional image generation task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as apache-2.0.
Paper: arXiv 2605.06137 · Code: github.com/Zyriix/prologue · Demo: 🤗 Space
Downloads · 30 days
0
Access
Public
Updated May 17, 2026
Repo size
68.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors34.4 GB · 50%
From the Hugging Face model README
Paper: arXiv 2605.06137 · Code: github.com/Zyriix/prologue · Demo: 🤗 Space
Bowen Zheng · Weijian Luo · Guang Yang · Colin Zhang · Tianyang Hu

Image =
[prologue tokens] + [visual tokens]. Prologue tokens are a small set of latent tokens prepended to the visual sequence and trained only with the AR cross-entropy loss. Visual tokens stay dedicated to reconstruction. The reconstruction–generation gap closes for free, and prologue tokens spontaneously develop semantic structure under pure CE gradients.
The figure above is generated with fixed prologue tokens, varying visual tokens only (Prologue-L–XL). Class identity and global layout stay; texture varies.
| Model | AR | rFID ↓ | gFID ↓ | gFID<sub>noCFG</sub> ↓ | IS ↑ | Pre. ↑ | Rec. ↑ |
|---|---|---|---|---|---|---|---|
| 1D Tokenizer (no CE) | 115M | 2.11 | 6.10 | 19.32 | — | — | — |
| 2D Tokenizer (no CE) | 115M | 2.15 | 5.02 | 21.01 | — | — | — |
| Prologue B–B | 115M | 2.24 | 4.11 | 10.75 | 210.3 | 0.83 | 0.48 |
| Prologue B–L | 305M | 2.24 | 2.67 | 6.56 | 251.2 | 0.82 | 0.56 |
| Prologue B–XL | 685M | 2.24 | 2.43 | 5.22 | 252.6 | 0.80 | 0.59 |
| Prologue L–B | 115M | 0.99 | 2.15 | 5.02 | 219.9 | 0.79 | 0.60 |
| Prologue L–L | 305M | 0.99 | 1.52 | 2.81 | 251.6 | 0.77 | 0.66 |
| Prologue L–XL | 685M | 0.99 | 1.46 | 2.26 | 257.7 | 0.78 | 0.66 |
| Prologue-Post (frozen 2D) | 115M | 2.15 | 3.88 | 11.04 | — | — | — |
| Prologue-OneStage (joint) | 115M | 2.09 | 5.41 | 21.00 | — | — | — |
All numbers are reproducible end-to-end with bash eval.sh in the GitHub repo. gFID / IS follow the ADM evaluation protocol (50k samples vs VIRTUAL_imagenet256_labeled.npz).
This Hugging Face Hub repository hosts all released weights for the paper — 6 tokenizers and 9 AR models (~63 GB total). The matching code, training scripts, and full documentation live on GitHub.
| Directory | rFID | Size | Note |
|---|---|---|---|
1d-tokenizer | 2.11 | 3.2 GB | 1D baseline, z_len = 256 |
2d-tokenizer | 2.15 | 3.2 GB | 2D baseline, z_len = 256 |
prologue-b-tokenizer | 2.24 | 4.1 GB | Prologue Base; VGG-LPIPS; codebook = 16 384 |
prologue-l-tokenizer | 0.99 | 6.7 GB | Prologue Large; ConvNeXt-logit; codebook = 4096; asymmetric decoder 24×1024 |
prologue-post-tokenizer | 2.15 | 3.2 GB | Prologue-Post (frozen 2D + new prologue path) |
prologue-onestage-joint | 2.09 | 5.5 GB | Joint AE + AR-Base, single-stage. AR shards live inside as model_5.safetensors / model_6.safetensors. |
| Directory | Pair with | Size | AR params | gFID (CFG) | gFID (no CFG) |
|---|---|---|---|---|---|
ar-1d-base | 1d-tokenizer | 1.8 GB | 115M | 6.10 | 19.32 |
ar-2d-base | 2d-tokenizer | 1.8 GB | 115M | 5.02 | 21.01 |
ar-prologue-b-b | prologue-b-tokenizer | 1.8 GB | 115M | 4.11 | 10.75 |
ar-prologue-b-l | prologue-b-tokenizer | 5.3 GB | 305M | 2.67 | 6.56 |
ar-prologue-b-xl | prologue-b-tokenizer | 11 GB | 685M | 2.43 | 5.22 |
ar-prologue-l-b | prologue-l-tokenizer | 1.5 GB | 115M | 2.15 | 5.02 |
ar-prologue-l-l | prologue-l-tokenizer | 4.9 GB | 305M | 1.52 | 2.81 |
ar-prologue-l-xl | prologue-l-tokenizer | 9.9 GB | 685M | 1.46 | 2.26 |
ar-prologue-post-b | prologue-post-tokenizer | 1.8 GB | 115M | 3.88 | 11.04 |
All checkpoints are safetensors following the 🤗 Accelerate convention (model.safetensors, model_1.safetensors, …). After download the layout matches the relative paths expected by eval.sh / app.py in the code repo — no mv step required.
pip install -U "huggingface_hub[cli]"
export HF_XET_HIGH_PERFORMANCE=1 # parallel Xet transfer
# everything (~63 GB)
hf download Zyriix/prologue --local-dir ckpts
# or just the headline model used by the demo (LXL, 9.9 GB + 6.7 GB tokenizer)
hf download Zyriix/prologue \
--include "ar-prologue-l-xl/*" \
--include "prologue-l-tokenizer/*" \
--local-dir ckpts
See the GitHub README for per-model commands and an inference-only slim layout (drops the ~50 % of bytes used for resuming training).
git clone https://github.com/Zyriix/prologue.git && cd prologue
bash setup_env.sh && conda activate prologue
# unpack the released ckpts (see above)
hf download Zyriix/prologue \
--include "ar-prologue-l-xl/*" \
--include "prologue-l-tokenizer/*" \
--local-dir ckpts
# (a) full headline-table reproduction
bash eval.sh
# (b) interactive Gradio demo: fix prologue, resample visual
python app.py
Programmatic loading inside Python:
from huggingface_hub import snapshot_download
ckpt_dir = snapshot_download(
repo_id="Zyriix/prologue",
allow_patterns=["ar-prologue-l-xl/*", "prologue-l-tokenizer/*"],
local_dir="ckpts",
max_workers=8,
)
# Then call into prologue/ as a library (load_models / sample_tokens):
# from sample_vis import load_models, sample_tokens
# See app.py for a full minimal example.
[prologue ; visual]; STE through the prologue codebook flows gradients into the encoder. Base: 150 epochs / Large: 200 epochs, both at batch size 256..npz.torch.compile, flash-attn 2.8 (source-built against torch 2.9.1 + cu128).NOTICE for the CC BY-NC-SA 4.0 carve-out on four NVIDIA StyleGAN3-derived files.@article{zheng2026prologue,
title = {Autoregressive Visual Generation Needs a Prologue},
author = {Zheng, Bowen and Luo, Weijian and Yang, Guang and Zhang, Colin and Hu, Tianyang},
journal = {arXiv preprint arXiv:2605.06137},
year = {2026},
url = {https://arxiv.org/abs/2605.06137}
}
Inspired by (chronological) LPIPS (2018) · vector-quantize-pytorch (2020) · VQGAN / taming-transformers (2020) · guided-diffusion (2021) · VAR (2024.04) · LlamaGen (2024.06) · TiTok (2024.06) · Open-MAGVIT2 (2024.09) · ImageFolder (2024.10) · AliTok (2025.06).