Downloads · 30 days
0
medmekk/deep-gemm
deep-gemm is a machine learning model from medmekk. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
DeepGEMM kernel for the Hugging Face kernel-builder infrastructure.
Downloads · 30 days
0
Access
Public
Updated Feb 16, 2026
Repo size
18.2 MB
Likes
0
Public
Click a slice to open those files.
.so18.2 MB · 94%
From the Hugging Face model README
DeepGEMM kernel for the Hugging Face kernel-builder infrastructure.
This package provides FP8/FP4/BF16 GEMM kernels, einsum, attention, and hyperconnection operations from DeepSeek-AI/DeepGEMM, adapted to the kernels-community build structure with torch library bindings.
pip install kernels
import kernels
kernels.install("kernels-community/DeepGEMM")
import deep_gemm
# FP8 GEMM: D = A @ B.T
deep_gemm.fp8_gemm_nt((a_fp8, sfa), (b_fp8, sfb), d)
# BF16 GEMM: D = A @ B.T
deep_gemm.bf16_gemm_nt(a_bf16, b_bf16, d)
# cuBLASLt GEMM
deep_gemm.cublaslt_gemm_nt(a, b, d)
DeepGEMM uses Just-In-Time (JIT) compilation for its CUDA kernels. The kernel
templates (.cuh files in include/deep_gemm/) are compiled at runtime using
NVCC or NVRTC. First invocations may be slower due to compilation; results are
cached in ~/.deep_gemm/ for subsequent calls.
The JIT-compiled kernels depend on CUTLASS headers (cute/, cutlass/) at
runtime. The package will automatically search for CUTLASS in these locations:
DG_CUTLASS_INCLUDE environment variable (direct path to include dir)CUTLASS_HOME environment variable ($CUTLASS_HOME/include)include/ directoryCUDA_HOME/include (some CUDA 12.8+ installs bundle cute/)nvidia-cutlass Python packageSet one of these if JIT compilation fails with missing CUTLASS headers:
export CUTLASS_HOME=/path/to/cutlass
# or
export DG_CUTLASS_INCLUDE=/path/to/cutlass/include