Downloads · 30 days
3
60% of all-time downloads
kernels-staging/vllm-flash-attn3
vllm-flash-attn3 is a machine learning model from kernels-staging. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for kernels. The card lists the license as apache-2.0.
This is the repository card of kernels-community/vllm-flash-attn3 that has been pushed on the Hub. It was built to be used with the kernels library. This card was automatically generated.
Downloads · 30 days
3
60% of all-time downloads
All-time downloads
5
Public
Repo size
1.6 GB
Likes
0
Public
Click a slice to open those files.
Other1.5 KB · 63%
From the Hugging Face model README
This is the repository card of kernels-community/vllm-flash-attn3 that has been pushed on the Hub. It was built to be used with the kernels library. This card was automatically generated.
# make sure `kernels` is installed: `pip install -U kernels`
from kernels import get_kernel
kernel_module = get_kernel("kernels-community/vllm-flash-attn3", version=1)
flash_attn_combine = kernel_module.flash_attn_combine
flash_attn_combine(...)
flash_attn_combineflash_attn_funcflash_attn_qkvpacked_funcflash_attn_varlen_funcflash_attn_with_kvcacheget_scheduler_metadataBenchmarking script is available for this kernel. Run kernels benchmark kernels-community/vllm-flash-attn3 --version 1.