Downloads · 30 days
6
15% of all-time downloads
pat883/attn-prefill-paged
attn-prefill-paged is a machine learning model from pat883. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for kernels.
Paged/chunked flash-attention prefill (+fp8) over a block-table KV cache.
Downloads · 30 days
6
15% of all-time downloads
All-time downloads
39
Public
Repo size
12.2 MB
Likes
0
Public
Click a slice to open those files.
.so4.1 MB · 100%
From the Hugging Face model README
Paged/chunked flash-attention prefill (+fp8) over a block-table KV cache.
Built with kernel-builder for AMD RDNA4 (gfx1201).
Load with the kernels library:
from kernels import get_kernel
kernel = get_kernel("pat883/attn-prefill-paged")
Requires a ROCm PyTorch build (torch 2.10 / ROCm 7.x) on an RDNA4 card. Built variants:
torch210-cxx11-rocm70andtorch210-cxx11-rocm71.