Downloads · 30 days
165
34% of all-time downloads
gsrunion/Ornith-1.0-9B-ROCmFPX-GGUF
Ornith-1.0-9B-ROCmFPX-GGUF is a text generation model from gsrunion. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as mit.
ROCmFPX-family GGUF quantizations of deepreinforce-ai/Ornith-1.0-9B, produced for AMD Strix Halo (Ryzen AI Max, gfx1151) and the hal0 home inference platform.
Downloads · 30 days
165
34% of all-time downloads
All-time downloads
488
Public
Repo size
28.8 GB
Likes
0
Public
Click a slice to open those files.
.gguf28.8 GB · 100%
From the Hugging Face model README
ROCmFPX-family GGUF quantizations of deepreinforce-ai/Ornith-1.0-9B, produced for AMD Strix Halo (Ryzen AI Max, gfx1151) and the hal0 home inference platform.
⚠️ These files require the Hal0ai/Hal0_ROCmFPX llama.cpp fork (or the
ghcr.io/hal0ai/hal0-rocmfpxcontainer image that hal0 uses). Stock llama.cpp will reject the tensor types (invalid ggml type 101).
| File | Quant | BPW | Size | Notes |
|---|---|---|---|---|
Ornith-1.0-9B-Q4_0_ROCMFP4_STRIX_LEAN.gguf | Q4_0_ROCMFP4_STRIX_LEAN | 4.42 | 4.96 GB | Size-biased Strix recipe, Q5_K token embeddings — best size/speed |
Ornith-1.0-9B-Q4_0_ROCMFP4_STRIX.gguf | Q4_0_ROCMFP4_STRIX | 4.54 | 5.09 GB | Quality-biased attention K/V recipe |
Ornith-1.0-9B-Q6_0_ROCMFPX_STRIX_QUALITY.gguf | Q6_0_ROCMFPX_STRIX_QUALITY | 7.50 | 8.41 GB | FP6 bulk + Q8 protected tensors — near-lossless daily driver |
Ornith-1.0-9B-Q8_0_ROCMFPX_AGENT.gguf | Q8_0_ROCMFPX_AGENT | 8.41 | 9.42 GB | Agent profile: protects embeddings, attn Q/K/V/O and select FFN tensors for tool-calling / JSON fidelity |
mmproj-BF16.gguf | BF16 | — | 0.92 GB | Vision projector (Ornith is multimodal) — load alongside any quant |
imatrix_unsloth.gguf_file | — | — | 5 MB | Importance matrix used for calibration (from unsloth, included for reproducibility) |
On AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB unified LPDDR5X, ROCm backend, hal0-rocmfpx image):
BF16 GGUF source and imatrix from unsloth/Ornith-1.0-9B-GGUF, quantized with the Hal0_ROCmFPX fork's llama-quantize:
llama-quantize --imatrix imatrix_unsloth.gguf_file \
Ornith-1.0-9B-BF16.gguf Ornith-1.0-9B-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
Q4_0_ROCMFP4_STRIX_LEAN
(same invocation per type for the other three)
hal0 model pull gsrunion/Ornith-1.0-9B-ROCmFPX-GGUF # or download + add-from-path
hal0 slot create ornith --type llm --hardware rocm --model <id> --ctx-size 32768
curl -sS --max-time 300 -X POST http://127.0.0.1:8080/api/slots/ornith/load