Downloads · 30 days
1.1K
42% of all-time downloads
zongzhuofan/EasyRef
EasyRef is a text-to-image model from zongzhuofan. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as apache-2.0.
Downloads · 30 days
1.1K
42% of all-time downloads
All-time downloads
2.6K
Public
Repo size
21.2 GB
Likes
3
Public
Click a slice to open those files.
.bin21.2 GB · 100%
From the Hugging Face model README
Project Page | Paper | Code | 🤗 Demo
</div>EasyRef is capable of modeling the consistent visual elements of various group image references with a single generalist multimodal LLM in a zero-shot setting.
<div align="center"> <img src='examples/framework.png'> </div>More visualization examples are available in our project page.
We provide the inference code of EasyRef with SDXL in easyref_demo.
scale=1.0 by default. Lowering the scale value leads to more diverse but less consistent generation results.If you find EasyRef useful for your research and applications, please cite us using this BibTeX:
@article{easyref,
title={EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM},
author={Zong, Zhuofan and Jiang, Dongzhi and Ma, Bingqi and Song, Guanglu and Shao, Hao and Shen, Dazhong and Liu, Yu and Li, Hongsheng},
journal={arXiv preprint arXiv:2412.09618},
year={2024}
}