Downloads · 30 days
0
rushabh14/TEMU-VTOFF
TEMU-VTOFF is a image-to-image model from rushabh14. Use it when you need one image transformed into another. It is set up for diffusers. The card lists the license as cc-by-nc-4.0.
<div align="center" <h1 align="center"TEMU-VTOFF</h1 <h3 align="center"Text-Enhanced MUlti-category Virtual Try-Off</h3 </div
Downloads · 30 days
0
Access
Public
Updated Jun 15, 2025
Parameters
2.1B
8.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.1 GB · 100%
From the Hugging Face model README
Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals Davide Lobba<sup>1,2,*</sup>, Fulvio Sanguigni<sup>2,3,*</sup>, Bin Ren<sup>1,2</sup>, Marcella Cornia<sup>3</sup>, Rita Cucchiara<sup>3</sup>, Nicu Sebe<sup>1</sup> <sup>1</sup>University of Trento, <sup>2</sup>University of Pisa, <sup>3</sup>University of Modena and Reggio Emilia <sup>*</sup> Equal contribution
</div> <div align="center"> <a href="https://arxiv.org/abs/2505.21062" style="margin: 0 2px;"> <img src="https://img.shields.io/badge/Paper-Arxiv_2505.21062-darkred.svg" alt="Paper"> </a> <a href="https://temu-vtoff-page.github.io/" style="margin: 0 2px;"> <img src='https://img.shields.io/badge/Webpage-Project-silver?style=flat&logo=&logoColor=orange' alt='Project Webpage'> </a> <a href="https://github.com/davidelobba/TEMU-VTOFF" style="margin: 0 2px;"> <img src="https://img.shields.io/badge/GitHub-Repo-blue.svg?logo=github" alt="GitHub Repository"> </a> <!-- The Hugging Face model badge will be automatically displayed on the model page --> </div>TEMU-VTOFF is a novel dual-DiT (Diffusion Transformer) architecture designed for the Virtual Try-Off task: generating in-shop images of garments worn by a person. By combining a pretrained feature extractor with a text-enhanced generation module, our method can handle occlusions, multiple garment categories, and ambiguous appearances. It further refines generation fidelity via a feature alignment module based on DINOv2.
This model is based on stabilityai/stable-diffusion-3-medium-diffusers. The uploaded weights correspond to the finetuned feature extractor and the VTOFF DiT module.
Our contribution can be summarized as follows: