Downloads · 30 days
213
36% of all-time downloads
Gorilla4X/Quacken-R1-14B-FP8
Quacken-R1-14B-FP8 is a text generation model from Gorilla4X. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as mit.
Native fp8 E4M3 GGUF of DeepSeek-R1-Distill-Qwen-14B (a reasoning model) for AMD RDNA4 (gfx1201 - Radeon AI PRO R9700 / RX 9070 / 9070 XT / W-series), quantized with AMD Quark from the full-precision BF16 weights by T…
Downloads · 30 days
213
36% of all-time downloads
All-time downloads
597
Public
Repo size
17.2 GB
Likes
0
Public
Click a slice to open those files.
.gguf17.2 GB · 100%
From the Hugging Face model README
Native fp8 E4M3 GGUF of DeepSeek-R1-Distill-Qwen-14B (a reasoning model) for AMD RDNA4 (gfx1201 - Radeon AI PRO R9700 / RX 9070 / 9070 XT / W-series), quantized with AMD Quark from the full-precision BF16 weights by The Rock8.
The Rock8's llama.cpp fork runs this fp8 on RDNA4's native WMMA fp8 tensor cores
(prefill) and v_dot4_f32_fp8_fp8 (decode) - not a dequant-to-f16 fallback.
F8E4M3), block-scaled, produced by AMD Quark from BF16.DeepSeek-R1-Distill-Qwen-14B-Quark-F8E4M3.gguf (16 GB).This is an R1-distill reasoning model. It emits a chain-of-thought inside
<think>...</think> before the final answer; llama.cpp / OpenAI-compatible servers
surface that as reasoning_content. To disable thinking for a turn, append
/no_think to the prompt (or set the chat template's thinking flag off). Expect
longer generations by default because of the reasoning trace.
| Metric | Value |
|---|---|
| Perplexity (wikitext, 20 chunks, n_ctx=512) | 8.97 |
Prefill pp512 | 2499 t/s |
Decode tg128 | 33.4 t/s |
Benched on a single R9700 (gfx1201).
# reasoning chat (keeps <think>)
llama-cli -m DeepSeek-R1-Distill-Qwen-14B-Quark-F8E4M3.gguf -ngl 99 \
-p "Solve step by step: a train travels 60 km in 40 minutes. What is its speed in km/h?"
# fast, no reasoning trace
llama-cli -m DeepSeek-R1-Distill-Qwen-14B-Quark-F8E4M3.gguf -ngl 99 -p "What do you call a dried grape? Answer in one word. /no_think"
# bench
llama-bench -m DeepSeek-R1-Distill-Qwen-14B-Quark-F8E4M3.gguf -ngl 99 -p 512 -n 128
podman run -d --rm --runtime crun --name lemonade \
--device /dev/kfd --device /dev/dri \
--group-add keep-groups --security-opt seccomp=unconfined \
-v /path/to/quacken-r1-14b:/models:ro \
-e MODEL=/models/DeepSeek-R1-Distill-Qwen-14B-Quark-F8E4M3.gguf -e MODEL_NAME=Quacken-R1-14B-FP8 \
-e HIP_VISIBLE_DEVICES=0 -p 13305:13305 \
ghcr.io/the-monk/the-rock8:rdna4-tr713 serve
Container (same image on each registry; --runtime crun is required for GPU):
ghcr.io/the-monk/the-rock8:rdna4-tr713 - docker.io/gorilla4x/the-rock8:rdna4-tr713 - quay.io/the-monk/the-rock8:rdna4-tr713
(images may not be pushed to every registry yet).
Every artifact links to the others - land on any one, reach them all.