Downloads · 30 days
0
adasdadsd/smallq-flash-attention-ascend
smallq-flash-attention-ascend is a machine learning model from adasdadsd. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
面向推测解码(Speculative Decoding)/ MTP / decode-time attention 场景的 AscendC 自定义算子。当 Query 序列长度很小(典型 1–8 token)、KV 长度很长时,本算子比通用 FlashAttention 更省 UB、更好并行。
Downloads · 30 days
0
Access
Public
Updated May 12, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.cpp19.5 KB · 38%
From the Hugging Face model README
面向推测解码(Speculative Decoding)/ MTP / decode-time attention 场景的 AscendC 自定义算子。当 Query 序列长度很小(典型 1–8 token)、KV 长度很长时,本算子比通用 FlashAttention 更省 UB、更好并行。
SmallqFlashAttention,aclnn API aclnnSmallqFlashAttention| 张量 | Shape | 说明 |
|---|---|---|
| Q (input) | [numHeads, qLen, headDim] | fp16 |
| K (input) | [numHeads, kvLen, headDim] | fp16 |
| V (input) | [numHeads, kvLen, headDim] | fp16 |
| O (output) | [numHeads, qLen, headDim] | fp16 |
O = softmax(Q · Kᵀ / √headDim) · V
| 参数 | 范围 |
|---|---|
| numHeads | 任意 ≥ 1 |
| qLen | 任意 ≥ 1(典型 1–8) |
| headDim | 任意 ≥ 1(无 16 对齐要求) |
| kvLen | 任意 ≥ 1 |
实测在 7 张 910B3 上并行验证了 256 用例 × 0 失败,覆盖:
hdPad = ceil(headDim/16)*16 对齐;输出回 GM 时压紧到 headDim stride,非 16 对齐尾部用 DataCopyPad UB→GM 字节级写出(硬件指令 copy_ubuf_to_gm_align_b16 使用字节级写使能,避免多 head 共享 cache line 的写竞争)smallq_flash_attention/ # 算子源码(独立可编译)
├── README.md
├── CMakeLists.txt
├── op_kernel/
│ ├── smallq_flash_attention.cpp
│ └── smallq_flash_attention_impl.h
└── op_host/
├── smallq_flash_attention_def.cpp
├── smallq_flash_attention_proto.cpp
├── smallq_flash_attention_tiling.h
└── smallq_flash_attention_tiling.cpp
tests/ # 验证程序
├── CMakeLists.txt
├── test_aclnn.cpp # aclnn 调用样例(单 case)
└── run_model_tests.py # 多卡并行模型 shape 测试驱动
将 smallq_flash_attention/ 目录放入 cann-recipes-infer 项目的 ops/ascendc/src/ 下,然后:
cd cann-recipes-infer/ops/ascendc
bash build.sh -n "smallq_flash_attention" -c "ascend910b"
# 部署(必须装到 opp 路径下,opp 优先级高于 vendors)
bash output/CANN-custom_ops-none-linux.aarch64.run \
--quiet --install-path=$ASCEND_HOME_PATH/opp/vendors/customize
yes | cp -rf $ASCEND_HOME_PATH/opp/vendors/customize/vendors/customize/. \
$ASCEND_HOME_PATH/opp/vendors/customize/
rm -rf $ASCEND_HOME_PATH/opp/vendors/customize/vendors
#include "aclnn_smallq_flash_attention.h"
aclTensor *qTensor, *kTensor, *vTensor, *oTensor;
// ... 创建 fp16 aclTensor,shape [numHeads, qLen|kvLen, headDim]
uint64_t workspaceSize = 0;
aclOpExecutor* executor = nullptr;
aclnnSmallqFlashAttentionGetWorkspaceSize(
qTensor, kTensor, vTensor, oTensor, &workspaceSize, &executor);
void* workspace = nullptr;
if (workspaceSize > 0) aclrtMalloc(&workspace, workspaceSize, ACL_MEM_MALLOC_HUGE_FIRST);
aclnnSmallqFlashAttention(workspace, workspaceSize, executor, stream);
aclrtSynchronizeStream(stream);
# 编译测试程序
cd tests
cmake -B build -DCANN_PATH=$ASCEND_HOME_PATH
cmake --build build -j
# 单 case
NUM_HEADS=32 Q_LEN=1 HEAD_DIM=128 KV_LEN=4096 DEVICE_ID=1 IO_DIR=/tmp/io \
./build/test_aclnn
# 多卡并行扫描真实模型 shape(默认 NPU 1–7)
python3 run_model_tests.py
Apache-2.0