Downloads · 30 days
17
12% of all-time downloads
JackBinary/Laguna-S-2.1-GGUF-ROCMFPX
Laguna-S-2.1-GGUF-ROCMFPX is a machine learning model from JackBinary. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as openmdw-1.1.
This is a quantization of unsloth/Laguna-S-2.1-GGUF (BF16) (the original model is Poolside's Laguna S 2.1, a 118B-total / 8B-activated MoE with 256 routed experts).
Downloads · 30 days
17
12% of all-time downloads
All-time downloads
145
Public
Repo size
64.6 GB
Likes
0
Public
Click a slice to open those files.
.gguf64.6 GB · 100%
From the Hugging Face model README
This is a quantization of unsloth/Laguna-S-2.1-GGUF (BF16)
(the original model is Poolside's Laguna S 2.1,
a 118B-total / 8B-activated MoE with 256 routed experts).
[!IMPORTANT] You need the ROCmFPX fork of llama.cpp (or a llama.cpp build with ROCmFPX support). This file uses the experimental
q4_0_rocmfp4_fast(type 101) andq8_0_rocmfpx(type 103) weight formats, which stock llama.cpp releases do not understand — loading it elsewhere will fail with an unknown tensor type error.
| Tensor group | Type | Count |
|---|---|---|
Routed experts: blk.N.ffn_{gate,up,down}_exps | q4_0_rocmfp4_fast (4.25 bpw) | 141 |
| Everything else quantizable (attention, shared experts, embeddings, output head) | q8_0_rocmfpx (8.25 bpw) | 386 |
| Norms, biases, router weights/scales | f32 (untouched) | 287 |
# from the ROCmFPX fork (CPU-only build works fine for quantization)
llama-quantize \
--tensor-type "ffn_(gate|up|down)_exps=q4_0_rocmfp4_fast" \
Laguna-S-2.1-BF16-00001-of-00005.gguf \
Laguna-S-2.1-Q8_0_ROCMFPX-Q4FAST-experts.gguf Q8_0_ROCMFPX
# then merged from 5 shards: llama-gguf-split --merge ...
Note the leading dense layer (blk.0) keeps its dense FFN at q8_0_rocmfpx — only the routed
expert tensors were overridden.
# build ROCmFPX for your GPU (see the repo README; e.g. Strix Halo):
env JOBS=16 scripts/build-strix-rocmfp4-mtp.sh
# run (Vulkan was fastest in upstream tests on Strix Halo):
./build-strix-rocmfp4/bin/llama-cli \
-m Laguna-S-2.1-Q8_0_ROCMFPX-Q4FAST-experts.gguf \
-dev Vulkan0 -ngl 999 -fa on --jinja
[!NOTE] Benchmarks are pending — placeholder table below.
| Backend / GPU | Prompt (tok/s) | Generation (tok/s) | Context | Notes |
|---|---|---|---|---|
| TBD | TBD | TBD | TBD | TBD |
Quality comparison vs BF16 source (perplexity / KLD): TBD.
unsloth/Laguna-S-2.1-GGUF (BF16 shards)openmdw-1.1 (inherited from the source model)