Downloads · 30 days
0
wangkanai/wan21-vae
wan21-vae is a text-to-video model from wangkanai. Use it when you need video from a text prompt. It is set up for diffusers. The card lists the license as other.
WAN2.1 VAE is a novel 3D causal Variational Autoencoder specifically designed for high-quality video generation and compression. This repository contains the standalone VAE component used in the WAN (Open and Advanced…
Downloads · 30 days
0
Access
Public
Updated Oct 14, 2025
Repo size
254 MB
Likes
5
Public
Click a slice to open those files.
.safetensors254 MB · 100%
From the Hugging Face model README
WAN2.1 VAE is a novel 3D causal Variational Autoencoder specifically designed for high-quality video generation and compression. This repository contains the standalone VAE component used in the WAN (Open and Advanced Large-Scale Video Generative Models) framework.
The WAN2.1 VAE represents a breakthrough in video compression and reconstruction technology, featuring:
E:\huggingface\wan21-vae\
└── vae/
└── wan/
└── wan21-vae.safetensors (243 MB)
| File | Size | Format | Description |
|---|---|---|---|
wan21-vae.safetensors | 243 MB | SafeTensors | WAN2.1 VAE weights |
Total Repository Size: 243 MB
import torch
from diffusers import AutoencoderKL
# Load the WAN2.1 VAE
vae = AutoencoderKL.from_pretrained(
"E:/huggingface/wan21-vae/vae/wan",
torch_dtype=torch.float16
).to("cuda")
print(f"VAE loaded: {vae.config}")
import torch
from diffusers import AutoencoderKL
from PIL import Image
import numpy as np
# Load VAE
vae = AutoencoderKL.from_pretrained(
"E:/huggingface/wan21-vae/vae/wan",
torch_dtype=torch.float16
).to("cuda")
# Prepare video frames (example with dummy data)
# Shape: [batch, channels, frames, height, width]
video_frames = torch.randn(1, 3, 16, 480, 720).half().to("cuda")
# Encode video to latent space
with torch.no_grad():
latents = vae.encode(video_frames).latent_dist.sample()
print(f"Latent shape: {latents.shape}")
print(f"Compression ratio: {np.prod(video_frames.shape) / np.prod(latents.shape):.2f}x")
import torch
from diffusers import AutoencoderKL
# Load VAE
vae = AutoencoderKL.from_pretrained(
"E:/huggingface/wan21-vae/vae/wan",
torch_dtype=torch.float16
).to("cuda")
# Decode latents back to video frames
# Assuming you have latents from encoding step
with torch.no_grad():
reconstructed_video = vae.decode(latents).sample
print(f"Reconstructed video shape: {reconstructed_video.shape}")
import torch
from diffusers import DiffusionPipeline, AutoencoderKL
# Load custom VAE
vae = AutoencoderKL.from_pretrained(
"E:/huggingface/wan21-vae/vae/wan",
torch_dtype=torch.float16
)
# Load WAN model with custom VAE
pipe = DiffusionPipeline.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B",
vae=vae,
torch_dtype=torch.float16
).to("cuda")
# Generate video
prompt = "A serene beach at sunset with waves crashing"
video = pipe(prompt, num_frames=16, height=480, width=720).frames
print(f"Generated video: {len(video)} frames")
# Use gradient checkpointing for lower memory usage
vae.enable_gradient_checkpointing()
# Use CPU offloading for very large videos
vae.enable_sequential_cpu_offload()
# Use attention slicing for reduced VRAM
vae.enable_attention_slicing(1)
# Compile model for faster inference (PyTorch 2.0+)
vae = torch.compile(vae, mode="reduce-overhead")
# Use xFormers for efficient attention
vae.enable_xformers_memory_efficient_attention()
# Use half precision for faster inference
vae = vae.half()
# Process multiple video clips efficiently
batch_size = 4
video_clips = torch.randn(batch_size, 3, 16, 480, 720).half().to("cuda")
with torch.no_grad():
latents = vae.encode(video_clips).latent_dist.sample()
This model is released under a custom WAN license. Please refer to the official WAN repository for detailed licensing terms and usage restrictions.
License Type: Other (Custom WAN License)
If you use this VAE in your research or applications, please cite the WAN project:
@misc{wan2025,
title={WAN: Open and Advanced Large-Scale Video Generative Models},
author={WAN-AI Team},
year={2025},
publisher={Hugging Face},
howpublished={https://huggingface.co/Wan-AI}
}
For questions, issues, or collaboration inquiries:
Version: v1.3 Last Updated: 2025-10-14 Model Size: 243 MB Format: SafeTensors