Downloads · 30 days
0
JackBinary/ROCmFPX-Quantize-Playbook
ROCmFPX-Quantize-Playbook is a machine learning model from JackBinary. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A field-tested, agent-ready playbook for producing ROCmFPX hybrid GGUF quantizations (q40rocmfp4fast / q80rocmfpx) with the ROCmFPX fork of llama.cpp — fully CPU-only, from either pre-quantized GGUF repos (e.g. Unslot…
Downloads · 30 days
0
Access
Public
Updated Aug 14, 2026
Repo size
—
Likes
1
Public
Click a slice to open those files.
.md13.4 KB · 90%
From the Hugging Face model README
A field-tested, agent-ready playbook for producing ROCmFPX hybrid GGUF quantizations
(q4_0_rocmfp4_fast / q8_0_rocmfpx) with the
ROCmFPX fork of llama.cpp — fully CPU-only,
from either pre-quantized GGUF repos (e.g. Unsloth BF16) or raw HF safetensors.
The complete guide is in rocmfpx-quantization-guide.md.
Hand this repo (or just that file) to a coding agent together with a Hugging Face link and a
recipe, and it can reproduce every step below.
llama-quantize / llama-gguf-split + Python conversion depsunsloth/*-GGUF)convert_hf_to_gguf.py → BF16 GGUF| Model | Recipe | Result |
|---|---|---|
| Laguna-S-2.1 118B-A10B | q4 experts / q8 rest | 224 GB → 61.6 GB (4.39 bpw) |
| Qwen3.8-27B | pure q8 and 16 GB hybrid from UD-Q4_K_XL tiers | 26.9 GB (8.25 bpw) / 16.4 GB (5.15 bpw) |
| G4-MeroMero-26B-A4B | q4 experts / q8 rest + mmproj | 50.5 GB → 14.0 GB (4.64 bpw) |
Timings on a 64-core CPU box: 118B MoE ≈ 8 min, 27B dense ≈ 1 min, 26B MoE ≈ 3 min.