Downloads · 30 days
0
wangkanai/wan22-fp8-i2v
wan22-fp8-i2v is a image-to-video model from wangkanai. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as other.
High-quality text-to-video (T2V) and image-to-video (I2V) generation models in FP8 quantized format for memory-efficient deployment on consumer-grade GPUs.
Downloads · 30 days
0
Access
Public
Updated Oct 27, 2025
Repo size
57.2 GB
Likes
2
Public
Click a slice to open those files.
.safetensors57.2 GB · 100%
From the Hugging Face model README
High-quality text-to-video (T2V) and image-to-video (I2V) generation models in FP8 quantized format for memory-efficient deployment on consumer-grade GPUs.
WAN 2.2 FP8 is a 14-billion parameter video generation model based on diffusion architecture, optimized with FP8 quantization for efficient deployment. This repository contains FP8 quantized variants that provide excellent quality with significantly reduced VRAM requirements compared to FP16 models (~50% memory reduction).
Key Features:
.safetensors formatModel Statistics:
.safetensors (secure tensor format)Located in diffusion_models/wan/
| Model | Size | Noise Schedule | Use Case |
|---|---|---|---|
wan22-t2v-14b-fp8-high-scaled.safetensors | 14GB | High-noise | Creative T2V, higher variance outputs |
wan22-t2v-14b-fp8-low-scaled.safetensors | 14GB | Low-noise | Faithful T2V, consistent results |
Total T2V models: 28GB
Located in diffusion_models/wan/
| Model | Size | Noise Schedule | Use Case |
|---|---|---|---|
wan22-i2v-14b-fp8-high-scaled.safetensors | 14GB | High-noise | Creative I2V, artistic interpretation |
wan22-i2v-14b-fp8-low-scaled.safetensors | 14GB | Low-noise | Faithful I2V, accurate reproduction |
Total I2V models: 26GB
| Model Type | Minimum VRAM | Recommended VRAM | GPU Examples |
|---|---|---|---|
| T2V FP8 | 16GB | 20GB+ | RTX 4080, RTX 3090, RTX 4070 Ti Super |
| I2V FP8 | 16GB | 20GB+ | RTX 4080, RTX 3090, RTX 4070 Ti Super |
System Requirements:
Compatible GPUs:
from diffusers import DiffusionPipeline
import torch
# Load T2V pipeline with FP8 support
pipe = DiffusionPipeline.from_pretrained(
"path-to-base-wan22-model",
torch_dtype=torch.float8_e4m3fn
)
# Load WAN 2.2 FP8 T2V model (low-noise for consistent results)
pipe.unet.from_single_file(
"E:/huggingface/wan22-fp8-i2v/diffusion_models/wan/wan22-t2v-14b-fp8-low-scaled.safetensors"
)
pipe.to("cuda")
# Generate video from text prompt
video = pipe(
prompt="a cat walking through a garden, cinematic, high quality",
num_inference_steps=50,
num_frames=16,
guidance_scale=7.5
).frames
# Save video
from diffusers.utils import export_to_video
export_to_video(video, "output_t2v.mp4", fps=8)
from diffusers import DiffusionPipeline
import torch
from PIL import Image
# Load input image
input_image = Image.open("path/to/your/image.jpg")
# Load I2V pipeline with FP8 support
pipe = DiffusionPipeline.from_pretrained(
"path-to-base-wan22-model",
torch_dtype=torch.float8_e4m3fn
)
# Load WAN 2.2 FP8 I2V model (high-noise for creative output)
pipe.unet.from_single_file(
"E:/huggingface/wan22-fp8-i2v/diffusion_models/wan/wan22-i2v-14b-fp8-high-scaled.safetensors"
)
pipe.to("cuda")
# Generate video from image
video = pipe(
image=input_image,
prompt="cinematic camera movement, high quality",
num_inference_steps=50,
num_frames=16,
guidance_scale=7.5
).frames
# Save video
from diffusers.utils import export_to_video
export_to_video(video, "output_i2v.mp4", fps=8)
# Enable memory optimizations for 16GB GPUs
pipe.enable_model_cpu_offload()
pipe.enable_xformers_memory_efficient_attention()
# Generate with reduced memory footprint
video = pipe(
prompt="your prompt here",
num_inference_steps=50,
num_frames=12, # Reduced from 16 for memory savings
guidance_scale=7.5
).frames
High-Noise Models (*-high-scaled.safetensors):
Low-Noise Models (*-low-scaled.safetensors):
Enable CPU Offloading: Offload model components to CPU when not in use
pipe.enable_model_cpu_offload()
Enable Attention Optimization: Use xformers for memory-efficient attention
pipe.enable_xformers_memory_efficient_attention()
Reduce Frame Count: Generate fewer frames for memory savings
num_frames=12 # Instead of 16
Sequential CPU Offload: Most aggressive memory savings
pipe.enable_sequential_cpu_offload()
Choose Appropriate Noise Schedule:
Increase Inference Steps: More steps = better quality (50-100 recommended)
num_inference_steps=75 # Higher quality, slower
Adjust Guidance Scale: Control prompt adherence (7.5 is standard)
guidance_scale=7.5 # Lower = more creative, Higher = more literal
num_inference_steps=30 # Faster, lower quality
| Content Type | Recommended Model | Reason |
|---|---|---|
| Realistic videos | Low-noise | Faithful reproduction, consistency |
| Artistic/abstract | High-noise | Creative interpretation, variety |
| Product demos | Low-noise | Predictable, professional results |
| Creative exploration | High-noise | Diverse outputs, experimentation |
| Production work | Low-noise | Consistent, reliable results |
| Task | Models | Description |
|---|---|---|
| Text-to-Video | wan22-t2v-* | Generate videos from text prompts only |
| Image-to-Video | wan22-i2v-* | Animate static images with text guidance |
"a cat walking through a garden, cinematic lighting, high quality, 4k"
"drone shot of mountain landscape at sunset, volumetric lighting"
"close-up of coffee being poured, slow motion, professional cinematography"
"time-lapse of city traffic at night, long exposure, urban photography"
"cinematic camera movement, smooth motion"
"gentle zoom in, professional cinematography"
"dynamic action, high energy movement"
"subtle animation, natural motion"
The model should NOT be used for:
Misuse Risks:
This repository uses the "other" license tag. Please check the original WAN 2.2 model repository for specific license terms, usage restrictions, and commercial use permissions.
If you use WAN 2.2 FP8 in your research or applications, please cite the original model:
@misc{wan22-fp8,
title={WAN 2.2 FP8: Text-to-Video and Image-to-Video Generation},
author={WAN Team},
year={2024},
howpublished={\url{https://huggingface.co/wan22}},
note={FP8 quantized variant}
}
Problem: CUDA out of memory during generation
Solutions:
pipe.enable_model_cpu_offload()pipe.enable_sequential_cpu_offload()num_frames=12 (instead of 16)pipe.enable_xformers_memory_efficient_attention()Problem: Generated videos have poor quality or artifacts
Solutions:
Problem: Video generation is too slow
Solutions:
pipe.enable_xformers_memory_efficient_attention()Problem: Cannot load model or incorrect format errors
Solutions:
from_single_file() method for safetensors loadingFor questions, issues, or contributions:
This model card was created following Hugging Face model card guidelines and best practices for responsible AI documentation.
Last Updated: October 14, 2025 Model Version: WAN 2.2 FP8 I2V v1.0 Repository Type: Quantized Model Weights Total Size: ~56GB (4 models × 14GB each)