Downloads · 30 days
0
lldacing/flash-attention-windows-wheel
flash-attention-windows-wheel is a machine learning model from lldacing. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as bsd-3-clause.
Downloads · 30 days
0
Access
Public
Updated May 31, 2025
Repo size
4.6 GB
Likes
333
Public
Click a slice to open those files.
.whl4.2 GB · 100%
From the Hugging Face model README
Windows wheels of flash-attention
Build cuda wheel steps
git clone https://github.com/Dao-AILab/flash-attention
cd flash-attention
v2.7.0.post2 (you can get latest tag by git describe --tags or list all available tags by git tag -l)git checkout -b v2.7.0.post2 v2.7.0.post2
Download WindowsWhlBuilder_cuda.bat into flash-attention
To build with MSVC, please open the "Native Tools Command Prompt for Visual Studio". The exact name may depend on your version of Windows, Visual Studio, and cpu architecture (in my case it was "x64 Native Tools Command Prompt for VS 2022".)
My Visual Studio Installer version

Switch python env and make sure the corresponding torch cuda version is installed
Start task
# Build with 1 parallel workers (I used 8 workers on i9-14900KF-3.20GHz-RAM64G, which took about 30 minutes.)
# If you want reset workers, you should edit `WindowsWhlBuilder_cuda.bat` and modify `set MAX_JOBS=1`.(I tried to modify it by parameters, but failed)
WindowsWhlBuilder_cuda.bat
# Build with sm80 and sm120
WindowsWhlBuilder_cuda.bat CUDA_ARCH="80;120"
# Enable cxx11abi
WindowsWhlBuilder_cuda.bat CUDA_ARCH="80;120" FORCE_CXX11_ABI=TRUE
dist directory