Downloads · 30 days
348
17% of all-time downloads
blowing-up-groundhogs/emuru_vae
emuru_vae is a machine learning model from blowing-up-groundhogs. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as mit.
Downloads · 30 days
348
17% of all-time downloads
All-time downloads
2.1K
Public
Parameters
14M
55.9 MB on disk
Likes
6
Public
Click a slice to open those files.
.safetensors55.9 MB · 100%
From the Hugging Face model README


This repository hosts the Emuru Convolutional VAE, described in our CVPR2025 paper. The model features a convolutional encoder and decoder, each with four layers. The output channels for these layers are 32, 64, 128, and 256, respectively. The encoder downsamples an input RGB image (with three channels and dimensions width and height) to a latent representation with a single channel and spatial dimensions that are one-eighth of the original height and width. This design compresses the style information in the image, allowing a lightweight Transformer Decoder to efficiently process the latent features. Training code is released on GitHub.
You can load the pre-trained Emuru VAE using Diffusers’ AutoencoderKL interface with a single line of code:
from diffusers import AutoencoderKL
model = AutoencoderKL.from_pretrained("blowing-up-groundhogs/emuru_vae")
Below is an example code snippet that demonstrates how to load an image directly from a URL, process it, encode it into the latent space, decode it back to image space, and save the reconstructed image.
import torch
from PIL import Image
from diffusers import AutoencoderKL
from huggingface_hub import hf_hub_download
from torchvision.transforms.functional import to_tensor, to_pil_image, normalize
# Load the pre-trained Emuru VAE from Hugging Face Hub.
model = AutoencoderKL.from_pretrained("blowing-up-groundhogs/emuru_vae")
# Function to preprocess an RGB image:
# Loads the image, converts it to RGB, and transforms it to a tensor normalized to [-1, 1].
def preprocess_image(image_path):
image = Image.open(image_path).convert("RGB")
image_tensor = to_tensor(image).unsqueeze(0) # Add batch dimension
image_tensor = normalize(image_tensor, [0.5], [0.5])
return image_tensor
# Function to postprocess a tensor back to a PIL image for visualization:
# Clamps the tensor to [-1, 1] and converts it to a PIL image.
def postprocess_tensor(tensor):
tensor = torch.clamp(tensor, -1, 1).squeeze(0) # Remove batch dimension
tensor = (tensor + 1) / 2
return to_pil_image(tensor)
# Example: Encode and decode an image.
# Replace the following line with your image path.
image_path = hf_hub_download(repo_id="blowing-up-groundhogs/emuru_vae", filename="samples/lam_sample.jpg")
input_image = preprocess_image(image_path)
# Encode the image to the latent space.
# The encode() method returns an object with a 'latent_dist' attribute.
# We sample from this distribution to obtain the latent representation.
with torch.no_grad():
latent_dist = model.encode(input_image).latent_dist
latents = latent_dist.sample()
# Decode the latent representation back to image space.
with torch.no_grad():
reconstructed = model.decode(latents).sample
# Load the original image for comparison.
original_image = Image.open(image_path).convert("RGB")
# Convert the reconstructed tensor back to a PIL image.
reconstructed_image = postprocess_tensor(reconstructed)
# Save the reconstructed image.
reconstructed_image.save("reconstructed_image.png")
If you use this VAE in your research or wish to refer to it, please cite:
@InProceedings{Pippi_2025_CVPR,
author = {Pippi, Vittorio and Quattrini, Fabio and Cascianelli, Silvia and Tonioni, Alessio and Cucchiara, Rita},
title = {Zero-Shot Styled Text Image Generation, but Make It Autoregressive},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2025},
pages = {7910-7919}
}