Downloads · 30 days
42
7% of all-time downloads
donghao-zhou/HiFi-Inpaint
HiFi-Inpaint is a image-to-image model from donghao-zhou. Use it when you need one image transformed into another. It is set up for diffusers. The card lists the license as cc-by-nc-4.0.
<h1 align="center" style="line-height: 50px;" HiFi-Inpaint: High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images </h1
Downloads · 30 days
42
7% of all-time downloads
All-time downloads
613
Public
Repo size
1.9 GB
Likes
2
Public
Click a slice to open those files.
.safetensors1.9 GB · 100%
From the Hugging Face model README
HiFi-Inpaint is a reference-based human-product image inpainting model for generating detail-preserving human-product images. Given a product reference image, a masked condition image, and a text prompt/caption, the model is designed to reconstruct the missing region while preserving fine-grained product appearance.
This repository contains the released model weights for HiFi-Inpaint, intended for research and model development on high-fidelity reference-guided inpainting.
HiFi-Inpaint/
├── README.md
├── alpha_blocks.pt
└── pytorch_lora_weights.safetensors
pytorch_lora_weights.safetensors: LoRA weights for the HiFi-Inpaint model.alpha_blocks.pt: auxiliary alpha-block weights used by the HiFi-Inpaint model pipeline.HiFi-Inpaint is intended for research and model development on reference-based human-product generation and inpainting. Typical use cases include:
This model is released as research weights and is not intended for deceptive, harmful, privacy-violating, or otherwise unlawful applications.
Please refer to the official code repository for installation, pipeline construction, and inference scripts:
https://github.com/Correr-Zhou/HiFi-Inpaint
A typical setup should download this repository's weights and load:
pytorch_lora_weights.safetensors as the LoRA checkpoint.alpha_blocks.pt as the auxiliary alpha-block checkpoint required by the inference pipeline.The model is associated with HP-Image-40K, a training dataset for high-fidelity reference-based human-product image inpainting. The dataset contains 43,632 aligned training samples with product reference images, ground-truth target images, masked condition images, binary masks, and captions.
Dataset repository: https://huggingface.co/datasets/donghao-zhou/HP-Image-40K
This model is released for research and model development purposes.
If you find this model useful in your research, please cite:
@article{liu2026hifiinpaint,
title={HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images},
author={Liu, Yichen and Zhou, Donghao and Wang, Jie and Gao, Xin and Liu, Guisheng and Li, Jiatong and Zhang, Quanwei and Lyu, Qiang and Guo, Lanqing and Wen, Shilei and Wang, Weiqiang and Heng, Pheng-Ann},
journal={arXiv preprint arXiv:2603.02210},
year={2026}
}
For questions about the model or dataset, please contact Donghao Zhou: [email protected].