Downloads · 30 days
0
Sumitc13/flash-attn-windows-wheels
flash-attn-windows-wheels is a machine learning model from Sumitc13. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for flash-attn. The card lists the license as bsd-3-clause.
Prebuilt flash-attn wheels for Windows, focused on combinations that aren't published anywhere else (newer PyTorch versions, CUDA 13, Python 3.12+, consumer Blackwell support).
Downloads · 30 days
0
Access
Public
Updated May 7, 2026
Repo size
240 MB
Likes
5
Trending 2
Click a slice to open those files.
.whl240 MB · 100%
From the Hugging Face model README
Prebuilt flash-attn wheels for Windows, focused on combinations that aren't published anywhere else (newer PyTorch versions, CUDA 13, Python 3.12+, consumer Blackwell support).
All wheels are built from the official Dao-AILab/flash-attention source and bundle kernels for multiple GPU architectures, so a single wheel works across most NVIDIA cards from Ampere through consumer Blackwell.
| flash-attn | CUDA | PyTorch | Python | ABI | File |
|---|---|---|---|---|---|
| 2.8.3 | 13.0 | 2.11.0 | 3.12 | True | flash_attn-2.8.3+cu130torch2.11.0cxx11abiTRUE-cp312-cp312-win_amd64.whl |
Want a combination that isn't here? Open a discussion on the Community tab.
Pick the wheel matching your environment and run:
pip install https://huggingface.co/Sumitc13/flash-attn-windows-wheels/resolve/main/<WHEEL_FILENAME>
Example:
pip install https://huggingface.co/Sumitc13/flash-attn-windows-wheels/resolve/main/flash_attn-2.8.3+cu130torch2.11.0cxx11abiTRUE-cp312-cp312-win_amd64.whl
The filename encodes everything you need to match:
flash_attn-{VERSION}+cu{CUDA}torch{TORCH}cxx11abi{ABI}-cp{PY}-cp{PY}-win_amd64.whl
└─2.8.3─┘ └─130─┘ └─2.11.0─┘ └─TRUE─┘ └─312─┘
Verify your environment first:
python -c "import torch; print(torch.__version__, torch.version.cuda)"
# e.g. "2.11.0+cu130 13.0" → use cu130 + torch2.11.0 wheel
Wheels are built with TORCH_CUDA_ARCH_LIST=8.0;8.6;8.9;9.0;12.0 unless noted otherwise, covering:
| Compute Capability | Architecture | Examples |
|---|---|---|
| 8.0 | Ampere | A100 |
| 8.6 | Ampere | RTX 30-series, A6000 |
| 8.9 | Ada Lovelace | RTX 40-series, L40S |
| 9.0 | Hopper | H100, H200 |
| 12.0 | Consumer Blackwell | RTX 5090, RTX Pro 6000 |
import torch
from flash_attn import flash_attn_func
q = torch.randn(1, 8, 128, 64, dtype=torch.float16, device='cuda')
print(flash_attn_func(q, q, q).shape)
# Expected: torch.Size([1, 8, 128, 64])
cu* tag in each wheel filenameNVCC_FLAGS=--compiler-options /Zc:preprocessor).BSD-3-Clause, matching upstream Flash Attention.
Unofficial community builds provided as-is with no warranty. Not affiliated with Dao-AILab or NVIDIA.