Downloads · 30 days
0
shanecol/llama.cpp-Strix-Halo-DFlash2-Windows-HIP
llama.cpp-Strix-Halo-DFlash2-Windows-HIP is a machine learning model from shanecol. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for llama.cpp. The card lists the license as mit.
This repository provides an unofficial, reproducible Windows x64 HIP build of llama-server.exe for AMD Strix Halo (gfx1151) with upstream DFlash/DFlash2 support and a one-file correction for silent long-prompt corrupt…
Downloads · 30 days
0
Access
Public
Updated Sep 7, 2026
Repo size
86.8 MB
Likes
0
Public
Click a slice to open those files.
.exe86.8 MB · 100%
From the Hugging Face model README
This repository provides an unofficial, reproducible Windows x64 HIP build of
llama-server.exe for AMD Strix Halo (gfx1151) with upstream DFlash/DFlash2
support and a one-file correction for silent long-prompt corruption.
It is not an official build from AMD, the llama.cpp maintainers, Hugging Face, or RUBI. It was built and validated on one Ryzen AI MAX+ 395 system. Read the validation boundary before treating it as production-ready.
Current llama.cpp supports draft-dflash, but its HIP integrated-GPU host-buffer
path can produce plausible-looking wrong logits on gfx1151 when a prompt is
split across micro-batches. The defect is documented in
ggml-org/llama.cpp#28211.
This build applies the narrowly scoped fix already present in public llama.cpp
commit 865374b: preserve integrated-device
classification while refusing direct HIP host-buffer compute specifically on
gfx1151.
bin/llama-server.exe6587aecb1f636954a6e1d787a896cc461fac343797a1fe9ab07a098ca011b3aa0.4.0-dev (build 1, commit 465e49b)465e49b9cea78a68b9c244ffb48d0ee24a82873d, tag b10830patches/gfx1151-disable-direct-host-buft.patchVerify before running:
Get-FileHash .\bin\llama-server.exe -Algorithm SHA256
The executable is not code-signed. A matching checksum proves it matches this upload; it does not substitute for reviewing and rebuilding the source.
gfx1151)PATHThe validated machine used:
32.0.31041.1004gfx1151-7.13.07.13.99004-3309c6114aROCm DLLs are deliberately not included. The executable dynamically imports
amdhip64_7.dll and hipblas.dll; install them through an appropriate TheRock
distribution and keep one coherent runtime version. See the
official llama.cpp Windows HIP guide.
The target and draft GGUF files are not included. DFlash drafts are trained for a specific target family; do not assume an unrelated draft is compatible merely because it loads.
.\scripts\run_dflash2_example.ps1 `
-ServerExe .\bin\llama-server.exe `
-TargetModel D:\Models\Target.gguf `
-DraftModel D:\Models\Target-DFlash2-Q8_0.gguf `
-RocmBin C:\TheRock\bin `
-Port 8080
The example uses p_min=0.90 and n_max=7. For a DFlash2 checkpoint with a
trained block size of eight, requesting n_max=8 is clamped by llama.cpp to seven
draft tokens, so seven states the effective setting directly.
Verify the live surface:
Invoke-RestMethod http://127.0.0.1:8080/health
Invoke-RestMethod http://127.0.0.1:8080/v1/models
(Invoke-WebRequest http://127.0.0.1:8080/metrics).Content
All rows below use the same Qwen3.8-27B UD-Q4_K_XL target weights. The old HIP
and fixed HIP rows used the locally available Q8 DFlash2 draft where identified.
| Runtime/configuration | Main-38 | Longctx + deep reasoning | Ladder, seed 1 |
|---|---|---|---|
| First-pass Windows HIP, DFlash2 Q8 | 0.7737 best recorded pass | 5/11 = 0.4545 | rung 8 pass; rung 9 miss |
| First-pass Windows HIP, self-MTP | 0.7842 best recorded pass | 5/11 = 0.4545 | rung 8 pass; rung 9 miss |
| Stock upstream b10588 Vulkan, self-MTP | 0.7158 final cap-adjusted pass | 11/11 = 1.0000 | rung 8 pass; rung 9 miss |
| Fixed b10830 Windows HIP, DFlash2 Q8 | not rerun | 11/11 = 1.0000 | rung 9/256 ops pass; rung 10/512 ops miss |
Additional fixed-build evidence:
ubatch=512 → PPL 17.1071; ubatch=2048 → PPL 17.1097 (0.015% difference).The public evidence/ directory contains the exact first-pass source
shims and a machine-readable, sanitized result summary. Raw private battery outputs
are deliberately excluded because publishing answer text would contaminate future
evaluation. The complete two-pass investigation is in
docs/DEBUGGING_WORKUP.md, and test interpretation is
in docs/VALIDATION.md.
The public build script takes explicit paths and refuses the wrong source commit:
git clone https://github.com/ggml-org/llama.cpp.git
git -C .\llama.cpp checkout 465e49b9cea78a68b9c244ffb48d0ee24a82873d
.\scripts\build_llama_windows_gfx1151fix.ps1 `
-SourceDir .\llama.cpp `
-ToolchainDir C:\TheRock `
-CMakeExe C:\Tools\CMake\bin\cmake.exe `
-NinjaExe C:\Tools\Ninja\ninja.exe `
-GitExe "C:\Program Files\Git\cmd\git.exe" `
-RcExe "C:\Program Files (x86)\Windows Kits\10\bin\10.0.26100.0\x64\rc.exe"
The validated build used CMake 4.4.3, Ninja 1.13.2, Git 2.55.0.windows.3,
Clang 23.0.0, and Windows SDK resource compiler 10.0.26100.0. Build flags are
documented in docs/BUILD_AND_RUNTIME.md.
Verified:
gfx1151 Radeon 8060S machineNot established:
--parallel > 1 production soak behaviorThe correction fixes a demonstrated HIP data-placement defect. It should not be described as improving model intelligence, DFlash acceptance, or every form of long-context behavior.
llama.cpp is MIT licensed; its license is included. The gfx1151 patch is derived
from the public llama.cpp commit linked above. The build, investigation, testing,
and documentation were produced collaboratively by Shane (shanecol), the RUBI
fleet, and OpenAI Codex. Public claims are summarized under evidence/; the raw
private battery records were retained by the project and are not redistributed.