Downloads Β· 30 days
855
6% of all-time downloads
htdong/Wan-Alpha_ComfyUI
Wan-Alpha_ComfyUI is a text-to-video model from htdong. Use it when you need video from a text prompt. It is set up for diffusers. The card lists the license as apache-2.0.
Downloads Β· 30 days
855
6% of all-time downloads
All-time downloads
13.7K
Public
Repo size
855 MB
Likes
31
Public
Click a slice to open those files.
.safetensors819 MB Β· 96%
From the Hugging Face model README
RGBA video generation, which includes an alpha channel to represent transparency, is gaining increasing attention across a wide range of applications. However, existing methods often neglect visual quality, limiting their practical usability. In this paper, we propose Wan-Alpha, a new framework that generates transparent videos by learning both RGB and alpha channels jointly. We design an effective variational autoencoder (VAE) that encodes the alpha channel into the RGB latent space. Then, to support the training of our diffusion transformer, we construct a high-quality and diverse RGBA video dataset. Compared with state-to-art methods, our model demonstrates superior performance in visual quality, motion realism, and transparency rendering. Notably, our model can generate a wide variety of semi-transparent objects, glowing effects, and fine-grained details such as hair strands. The released model is available on our website: this https URL .
<div align="center"> <h1> Wan-Alpha </h1> <h3>Wan-Alpha: High-Quality Text-to-Video Generation with Alpha Channel</h3> </div> <img src="assets/teaser.png" alt="Wan-Alpha Qualitative Results" style="max-width: 100%; height: auto;">Qualitative results of video generation using Wan-Alpha. Our model successfully generates various scenes with accurate and clearly rendered transparency. Notably, it can synthesize diverse semi-transparent objects, glowing effects, and fine-grained details such as hair.
| Prompt | Preview Video | Alpha Video |
|---|---|---|
| "Medium shot. A little girl holds a bubble wand and blows out colorful bubbles that float and pop in the air. The background of this video is transparent. Realistic style." | <img src="assets/girl.gif" width="320" height="180" style="object-fit:contain; display:block; margin:auto;"/> | <img src="assets/girl_pha.gif" width="335" height="180" style="object-fit:contain; display:block; margin:auto;"/> |
# Clone the project repository
git clone https://github.com/WeChatCV/Wan-Alpha.git
cd Wan-Alpha
# Create and activate Conda environment
conda create -n Wan-Alpha python=3.11 -y
conda activate Wan-Alpha
# Install dependencies
pip install -r requirements.txt
Download Wan2.1-T2V-14B
Download Lightx2v-T2V-14B
Download Wan-Alpha VAE
You can test our model through:
torchrun --nproc_per_node=8 --master_port=29501 generate_dora_lightx2v.py --size 832*480\
--ckpt_dir "path/to/your/Wan-2.1/Wan2.1-T2V-14B" \
--dit_fsdp --t5_fsdp --ulysses_size 8 \
--vae_lora_checkpoint "path/to/your/decoder.bin" \
--lora_path "path/to/your/epoch-13-1500.safetensors" \
--lightx2v_path "path/to/your/lightx2v_T2V_14B_cfg_step_distill_v2_lora_rank64_bf16.safetensors" \
--sample_guide_scale 1.0 \
--frame_num 81 \
--sample_steps 4 \
--lora_ratio 1.0 \
--lora_prefix "" \
--prompt_file ./data/prompt.txt \
--output_dir ./output
You can specify the weights of Wan2.1-T2V-14B with --ckpt_dir, LightX2V-T2V-14B with --lightx2v_path, Wan-Alpha-VAE with --vae_lora_checkpoint, and Wan-Alpha-T2V with --lora_path. Finally, you can find the rendered RGBA videos with a checkerboard background and PNG frames at --output_dir.
Prompt Writing Tip: You need to specify that the background of the video is transparent, the visual style, the shot type (such as close-up, medium shot, wide shot, or extreme close-up), and a description of the main subject. Prompts support both Chinese and English input.
# An example of prompt.
This video has a transparent background. Close-up shot. A colorful parrot flying. Realistic style.
Note: We have reorganized our models to ensure they can be easily loaded into ComfyUI. Please note that these models differ from the ones mentioned above.
ComfyUI/models folder and organize them as follows:ComfyUI/models
βββ diffusion_models
β βββ wan2.1_t2v_14B_fp16.safetensors
βββ loras
β βββ epoch-13-1500_changed.safetensors
β βββ lightx2v_T2V_14B_cfg_step_distill_v2_lora_rank64_bf16.safetensors
βββ text_encoders
β βββ umt5_xxl_fp8_e4m3fn_scaled.safetensors
βββ vae
β βββ wan_alpha_2.1_vae_alpha_channel.safetensors.safetensors
β βββ wan_alpha_2.1_vae_rgb_channel.safetensors.safetensors
ComfyUI/custom_nodes folder.This project is built upon the following excellent open-source projects:
We sincerely thank the authors and contributors of these projects.
If you find our work helpful for your research, please consider citing our paper:
@misc{dong2025wanalpha,
title={Wan-Alpha: High-Quality Text-to-Video Generation with Alpha Channel},
author={Haotian Dong and Wenjing Wang and Chen Li and Di Lin},
year={2025},
eprint={2509.24979},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2509.24979},
}
If you have any questions or suggestions, feel free to reach out via GitHub Issues . We look forward to your feedback!