Downloads · 30 days
277
64% of all-time downloads
kootaro/Wan2.2-Custom-Models-GGUF
Wan2.2-Custom-Models-GGUF is a image-to-video model from kootaro. Use it for the image-to-video task on the model card, and read the license before you ship it in a product. It is set up for gguf. The card lists the license as apache-2.0.
This repository provides highly optimized Wan2.2 Image-to-Video (I2V) GGUF and specialized custom models. These variants are tailored for running efficiently on memory-constrained environments, such as Google Colab eq…
Downloads · 30 days
277
64% of all-time downloads
All-time downloads
433
Public
Repo size
400 GB
Likes
0
Public
Click a slice to open those files.
.safetensors242 GB · 61%
From the Hugging Face model README
This repository provides highly optimized Wan2.2 Image-to-Video (I2V) GGUF and specialized custom models. These variants are tailored for running efficiently on memory-constrained environments, such as Google Colab equipped with an NVIDIA Tesla T4 GPU, while offering professional-grade motion extensions.
umt5_xxl_fp16 - umt5_xxl_fp8_e4m3fn_scaled), VAEs (Wan2_1_VAE_fp32 / wan_2.1_bf16), and specific FP8 integrated models are fully active and optimized for deployment. If you require raw unquantized BF16 weights, please wait for future repository syncs or utilize the available GGUF variants.To achieve perfect video motion without artifacts or image degradation (preventing fried, burnt, or oversaturated visuals), we strongly recommend using the following parameters:
| Parameter | Recommended Value | Note |
|---|---|---|
| Total Sampling Steps | 4 - 12 | Absolute maximum ceiling is 12 total steps for Lightning / Distilled V2 |
| CFG Scale | 1.0 - 2 | Crucial for preventing burnt images |
| High Noise Steps | 2, 4, 6, or 8 | To lock in strong motion. Can be split evenly (e.g., 8 steps total = 4 High / 4 Low) or 6 step total = 3 High / 3 low |
| Low Noise Steps | Dynamic (End Step: 4 - 12) | CRITICAL: The target End Step for Low Noise must NEVER exceed the Total Sampling Steps! |
| Sampler / Scheduler | euler + simple | Standard diffusion setup (Optionally, uni_pc can also be used for alternative fast-stepping) |
If you want to achieve higher visual fidelity and enhance micro-details, adopting a hybrid multi-pass approach is highly recommended. This strategy significantly sharpens fine details, effectively eliminates motion blur, and prevents fried visuals.
However, due to severe hardware VRAM limitations and Web GUI overhead, you MUST strictly adhere to the following setup configurations based on your execution environment:
Q8_H (High Noise) + Q8_H (Low Noise) GGUF files.Q8_H (High Noise) and chain it with wan2.2_i2v_low_noise_14B_fp8_scaled.safetensors as the final step to achieve ultimate sharpness and micro-details./diffusion_models folder, nor any external FP8 models placed outside in the root directory! Because the Web GUI consumes a massive amount of VRAM just to render its interface, available memory is extremely critical. Forcing these models via GUI will trigger an immediate OOM (Out of Memory) crash.
Q4_K_M to Q8_H range for both High Noise and Low Noise GGUF models.Q4K_M (High Noise) + Q4K_M (Low Noise).Q4_K (High Noise) + Q6_K (Low Noise) or Q4K_M.gguf (High Noise) + Q4K_M (Low Noise). This delivers optimized speed while maintaining excellent visual quality compared to full high-quants.Choose the right variant based on your creative workflow and VRAM configuration. All files are organized into dedicated subdirectories for pipeline flexibility:
/diffusion_models)These models feature pre-baked pipelines integrated with SVI (Stable Video Infinity) for continuous video synthesis and Consistent Face weights to prevent character distortion across frames.
Wan2_2-I2V-A14B-HIGH_SVI_consistent_face_nsfw_fp8.safetensors: Structural expert optimized for initial motion pathways, camera dynamics, and uncensored/free-form pipeline generations.Wan2_2-I2V-A14B-LOW_SVI_consistent_face_nsfw_fp8.safetensors: Fine-tuning expert optimized for character preservation, facial structural lock, and detailed refinement.wan2.2_i2v_high_noise_14B_...): Best for creative, high-motion generation, and diverse camera movements. Available in: Q4_K_M, Q6_K_L, Q6_K, Q8_H, and fp8_scaled.wan2.2_i2v_low_noise_14B_...): Best for high fidelity, generation stability, and strictly adhering to the prompt or structural layout of your starting frame. Available in: Q4_K_M, Q6_K_L, Q6_K, Q8_H, and fp8_scaled./loras: Contains raw targeted weights (high_noise and low_noise rank64 lightx2v 4-step) for modular multi-pass setups./text_encoders: Contains umt5_xxl_fp16.safetensors (11.4 GB) to maximize text prompt processing accuracy./vae: Contains Wan2_1_VAE_fp32.safetensors (508 MB) to prevent color degradation and artifacting during final video decoding.wan_2.1_bf16_vae.safetensors is also placed in the root directory for extra VRAM safety during low-tier runs.