Skip to content

sdfafdfsdf

vllm-flash-attn3

sdfafdfsdf/vllm-flash-attn3

vllm-flash-attn3 is a machine learning model from sdfafdfsdf. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for kernels. The card lists the license as apache-2.0.

This is an implementation of Flash Attention 3 CUDA kernels with support for attention sinks. The attention sinks implementation was contributed to Flash Attention by the vLLM team. The transformers team packaged the…

Downloads · 30 days

12

21% of all-time downloads

All-time downloads

58

Public

Repo size

16.2 GB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.so10.7 GB · 100%

At a glance

Library
kernels
License
apache-2.0
Access
Public
Created
Feb 20, 2026
Updated
Feb 20, 2026
SHA
89ba29ba
Library
kernels
License
apache-2.0
Created
Feb 20, 2026
Updated
Feb 20, 2026