Downloads · 30 days
228
100% of all-time downloads
bytedance-research/USO
USO is a text-to-image model from bytedance-research. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as apache-2.0.
<p align="center" <img src="assets/uso.webp" width="100"/ <p <h3 align="center" Unified Style and Subject-Driven Generation via Disentangled and Reward Learning </h3
Downloads · 30 days
228
100% of all-time downloads
All-time downloads
228
Public
Repo size
501 MB
Likes
191
Public
Click a slice to open those files.
.safetensors500 MB · 100%
From the Hugging Face model README
Paper: USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
<p align="center"> <a href="https://github.com/bytedance/USO"><img alt="Build" src="https://img.shields.io/github/stars/bytedance/USO"></a> <a href="https://bytedance.github.io/USO/"><img alt="Build" src="https://img.shields.io/badge/Project%20Page-USO-blue"></a> <a href="https://arxiv.org/abs/2508.18966"><img alt="Build" src="https://img.shields.io/badge/Tech%20Report-USO-b31b1b.svg"></a> <a href="https://huggingface.co/bytedance-research/USO"><img src="https://img.shields.io/static/v1?label=%F0%9F%A4%97%20Hugging%20Face&message=Model&color=green"></a> </p>
Existing literature typically treats style-driven and subject-driven generation as two disjoint tasks: the former prioritizes stylistic similarity, whereas the latter insists on subject consistency, resulting in an apparent antagonism. We argue that both objectives can be unified under a single framework because they ultimately concern the disentanglement and re-composition of content and style, a long-standing theme in style-driven research. To this end, we present USO, a Unified Style-Subject Optimized customization model. First, we construct a large-scale triplet dataset consisting of content images, style images, and their corresponding stylized content images. Second, we introduce a disentangled learning scheme that simultaneously aligns style features and disentangles content from style through two complementary objectives, style-alignment training and content-style disentanglement training. Third, we incorporate a style reward-learning paradigm denoted as SRL to further enhance the model's performance. Finally, we release USO-Bench, the first benchmark that jointly evaluates style similarity and subject fidelity across multiple metrics. Extensive experiments demonstrate that USO achieves state-of-the-art performance among open-source models along both dimensions of subject consistency and style similarity. Code and model: this https URL
Install the requirements
## create a virtual environment with python >= 3.10 <= 3.12, like
python -m venv uso_env
source uso_env/bin/activate
## or
conda create -n uso_env python=3.10 -y
conda activate uso_env
## then install the requirements by you need
pip install -r requirements.txt # legacy installation command
Then download checkpoints in one of the following ways:
# 1. download USO official checkpoints
pip install huggingface_hub
huggingface-cli download bytedance-research/USO --local-dir <YOU_SAVE_DIR> --local-dir-use-symlinks False
# 2. Then set the environment variable for FLUX.1 base model
export AE="YOUR_AE_PATH"
export FLUX_DEV="YOUR_FLUX_DEV_PATH"
export T5="YOUR_T5_PATH"
export CLIP="YOUR_CLIP_PATH"
# or export HF_HOME="YOUR_HF_HOME"
# 3. Then set the environment variable for USO
export LORA="<YOU_SAVE_DIR>/uso_flux_v1.0/dit_lora.safetensors"
export PROJECTION_MODEL="<YOU_SAVE_DIR>/uso_flux_v1.0/projector.safetensors"
hf_hub_download function in the code.Start from the examples below to explore and spark your creativity. ✨
# the first image is a content reference, and the rest are style references.
# for subject-driven generation
python inference.py --prompt "The man in flower shops carefully match bouquets, conveying beautiful emotions and blessings with flowers. " --image_paths "assets/gradio_examples/identity1.jpg" --width 1024 --height 1024
# for style-driven generation
# please keep the first image path empty
python inference.py --prompt "A cat sleeping on a chair." --image_paths "" "assets/gradio_examples/style1.webp" --width 1024 --height 1024
# for ip-style generation
python inference.py --prompt "The woman gave an impassioned speech on the podium." --image_paths "assets/gradio_examples/identity2.webp" "assets/gradio_examples/style2.webp" --width 1024 --height 1024
# for multi-style generation
# please keep the first image path empty
python inference.py --prompt "A handsome man." --image_paths "" "assets/gradio_examples/style3.webp" "assets/gradio_examples/style4.webp" --width 1024 --height 1024
We also appreciate it if you could give a star ⭐ to our Github repository. Thanks a lot!
If you find this project useful for your research, please consider citing our paper:
@article{wu2025uso,
title={USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning},
author={Shaojin Wu and Mengqi Huang and Yufeng Cheng and Wenxu Wu and Jiahe Tian and Yiming Luo and Fei Ding and Qian He},
year={2025},
eprint={2508.18966},
archivePrefix={arXiv},
primaryClass={cs.CV},
}