Downloads · 30 days
0
bytedance-research/UNO
UNO is a image-to-image model from bytedance-research. Use it when you need one image transformed into another. It is set up for transformers. The card lists the license as apache-2.0.
<h3 align="center" Less-to-More Generalization: Unlocking More Controllability by In-Context Generation </h3
Downloads · 30 days
0
Access
Public
Updated Aug 22, 2025
Repo size
1.9 GB
Likes
182
Public
Click a slice to open those files.
.safetensors1.9 GB · 100%
From the Hugging Face model README
<p align="center"> <span style="color:#137cf3; font-family: Gill Sans">Shaojin Wu,</span><sup></sup></a> <span style="color:#137cf3; font-family: Gill Sans">Mengqi Huang</span><sup>*</sup>,</a> <span style="color:#137cf3; font-family: Gill Sans">Wenxu Wu,</span><sup></sup></a> <span style="color:#137cf3; font-family: Gill Sans">Yufeng Cheng,</span><sup></sup> </a> <span style="color:#137cf3; font-family: Gill Sans">Fei Ding</span><sup>+</sup>,</a> <span style="color:#137cf3; font-family: Gill Sans">Qian He</span></a> <br> <span style="font-size: 16px">Intelligent Creation Team, ByteDance</span></p>

In this study, we propose a highly-consistent data synthesis pipeline to tackle this challenge. This pipeline harnesses the intrinsic in-context generation capabilities of diffusion transformers and generates high-consistency multi-subject paired data. Additionally, we introduce UNO, which consists of progressive cross-modal alignment and universal rotary position embedding. It is a multi-image conditioned subject-to-image model iteratively trained from a text-to-image model. Extensive experiments show that our method can achieve high consistency while ensuring controllability in both single-subject and multi-subject driven generation.
Clone our Github repo
Install the requirements
## create a virtual environment with python >= 3.10 <= 3.12, like
# python -m venv uno_env
# source uno_env/bin/activate
# then install
pip install -r requirements.txt
then download checkpoints in one of the three ways:
hf_hub_download function in the code to your $HF_HOME(the default value is ~/.cache/huggingface).huggingface-cli download <repo name> to download black-forest-labs/FLUX.1-dev, xlabs-ai/xflux_text_encoders, openai/clip-vit-large-patch14, TODO UNO hf model, then run the inference scripts.huggingface-cli download <repo name> --local-dir <LOCAL_DIR> to download all the checkpoints menthioned in 2. to the directories your want. Then set the environment variable TODO. Finally, run the inference scripts.python app.py
dreambench to download the dataset.git submodule update --init
python inference.py
accelerate launch train.py

For the purpose of fostering research and the open-source community, we plan to open-source the entire project, encompassing training, inference, weights, etc. Thank you for your patience and support! 🌟
If UNO is helpful, please help to ⭐ the repo.
If you find this project useful for your research, please consider citing our paper:
@article{wu2025less,
title={Less-to-More Generalization: Unlocking More Controllability by In-Context Generation},
author={Wu, Shaojin and Huang, Mengqi and Wu, Wenxu and Cheng, Yufeng and Ding, Fei and He, Qian},
journal={arXiv preprint arXiv:2504.02160},
year={2025}
}