Downloads · 30 days
0
pseudobacon/SageAttention2
SageAttention2 is a machine learning model from pseudobacon. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Pre-built SageAttention 2.2.0 for NVIDIA RTX 5090 (sm120 Blackwell) on Linux. INT8 quantized attention achieving 2-3x speedup over FlashAttention2 without accuracy loss.
Downloads · 30 days
0
Access
Public
Updated May 16, 2026
Repo size
5.7 MB
Likes
0
Public
Click a slice to open those files.
.whl5.7 MB · 100%
From the Hugging Face model README
Pre-built SageAttention 2.2.0 for NVIDIA RTX 5090 (sm_120 Blackwell) on Linux. INT8 quantized attention achieving 2-3x speedup over FlashAttention2 without accuracy loss.
Built from woct0rdho/SageAttention — a fork of thu-ml/SageAttention with ABI3 stable wheel support and additional fixes.
sageattention-2.2.0+cu130torch2.11.0andhigher-cp39-abi3-linux_x86_64.whl
| Component | Version |
|---|---|
| GPU | NVIDIA RTX 50 series (sm_120 Blackwell) |
| OS | Linux x86_64 |
| Python | 3.9+ (ABI3 stable — compatible with Python 3.9 or higher) |
| PyTorch | 2.11.0 or higher |
| CUDA | 13.0 or higher |
| Triton | compatible with your PyTorch version |
pip install sageattention-2.2.0+cu130torch2.11.0andhigher-cp39-abi3-linux_x86_64.whl --no-deps
Important: Always use
--no-depsto prevent pip from overwriting your PyTorch installation.
import torch
print(torch.__version__)
from sageattention import sageattn_varlen
print("SageAttention 2.2.0 OK")
Apache 2.0