Downloads · 30 days
273
12% of all-time downloads
charles2530/Wan2.2-NVFP4-Sparse
Wan2.2-NVFP4-Sparse is a machine learning model from charles2530. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as apache-2.0.
An extremely efficient Wan 2.2 14B variant: NVFP4 Quantization-Aware Step Distillation with Sparse Attention for Blackwell Architecture
Downloads · 30 days
273
12% of all-time downloads
All-time downloads
2.4K
Public
Repo size
67 GB
Likes
0
Public
Click a slice to open those files.
.safetensors67 GB · 100%
From the Hugging Face model README
An extremely efficient Wan 2.2 14B variant: NVFP4 Quantization-Aware Step Distillation with Sparse Attention for Blackwell Architecture
We strongly recommend using the official LightX2V Docker image for the cleanest environment and best reproducibility.
# 1. Pull LightX2V Docker image
docker pull lightx2v/lightx2v:26052801-cu130-5090
# 2. Run inference
bash scripts/wan22/distill/run_wan22_moe_t2v_extreme.sh
If Docker is not available, install the environment manually:
# 1. Install LightX2V
git clone https://github.com/ModelTC/LightX2V.git
cd LightX2V
uv pip install -v .
# 2. Install NVFP4 Kernel
pip install scikit_build_core uv
git clone https://github.com/NVIDIA/cutlass.git
cd lightx2v_kernel
MAX_JOBS=$(nproc) CMAKE_BUILD_PARALLEL_LEVEL=$(nproc) \
uv build --wheel \
-Cbuild-dir=build . \
-Ccmake.define.CUTLASS_PATH=/path/to/cutlass \
--verbose --color=always --no-build-isolation
pip install dist/*whl --force-reinstall --no-deps
# 3. Run inference
bash scripts/wan22/distill/run_wan22_moe_t2v_extreme.sh
Script: run_wan22_moe_t2v_extreme.sh
| Resolution | Wan2.2-T2V-14B | Wan2.2-NVFP4-Sparse |
|---|---|---|
| 480p | <video controls style="width: 260px; height: 180px; border-radius: 6px; object-fit: cover;" src="https://cdn-uploads.huggingface.co/production/uploads/658e760cccbc1e2cc78b4258/WTHhrzx7XR4S1Ys_6Kzx4.mp4"></video> | <video controls style="width: 260px; height: 180px; border-radius: 6px; object-fit: cover;" src="https://cdn-uploads.huggingface.co/production/uploads/658e760cccbc1e2cc78b4258/zorpw7gm9At0J2kCmvkDr.mp4"></video> |
| 720p | <video controls style="width: 260px; height: 180px; border-radius: 6px; object-fit: cover;" src="https://cdn-uploads.huggingface.co/production/uploads/658e760cccbc1e2cc78b4258/vkiyKj7CJA-r0yTz7TEum.mp4"></video> | <video controls style="width: 260px; height: 180px; border-radius: 6px; object-fit: cover;" src="https://cdn-uploads.huggingface.co/production/uploads/658e760cccbc1e2cc78b4258/TuECbzvW5jI9NHG6GLvIR.mp4"></video> |
Test Environment: RTX 5090 Single GPU | LightX2V Framework | End-to-End Latency
| Resolution | Wan2.2-T2V-14B | Wan2.2-NVFP4-Sparse | Speedup |
|---|---|---|---|
| 480p | 734s | 14.15s | 51.9x |
| 720p | 2668s | 45s | 59.3x |
lightx2v/lightx2v:26052801-cu130-5090.If you find this project helpful, please give us a ⭐ on GitHub
For questions or issues, please open an issue on LightX2V or contact [email protected].
</div>