Downloads · 30 days
0
ajokela/llama-cpp-kernels-gfx1151
llama-cpp-kernels-gfx1151 is a machine learning model from ajokela. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Pre-compiled HIP/ROCm kernels for llama-cpp-python on AMD gfx1151 GPUs (RDNA 4).
Downloads · 30 days
0
Access
Public
Updated Nov 27, 2025
Repo size
30.4 MB
Likes
0
Public
Click a slice to open those files.
.gz30.4 MB · 100%
From the Hugging Face model README
Pre-compiled HIP/ROCm kernels for llama-cpp-python on AMD gfx1151 GPUs (RDNA 4).
comgr/ - Precompiled LLVM kernel cache (~94MB)libggml-hip.so - HIP-accelerated GGML library compiled for gfx1151libggml*.so - Supporting GGML librarieslibllama.so - llama.cpp librarylibmtmd.so - Multi-modal libraryCMAKE_ARGS="-DGGML_HIPBLAS=on" AMDGPU_TARGETS="gfx1151" pip install llama-cpp-python
# Copy comgr cache (saves ~50 mins of JIT compilation!)
cp -r comgr ~/.cache/
# Optionally replace the compiled libs (must match llama-cpp-python version)
# cp lib*.so ~/.local/lib/python3.x/site-packages/llama_cpp/lib/
Required environment variables for gfx1151:
export HSA_OVERRIDE_GFX_VERSION=11.5.1
export AMDGPU_TARGETS=gfx1151
First-time loading of large GGML models on gfx1151 requires JIT compilation of GPU kernels, which can take 30-60+ minutes. This cache contains pre-compiled kernels that skip that initial compilation step.
MIT