Downloads · 30 days
62
27% of all-time downloads
Michionlion/Astrea-R8-Chat-9B-GGUF
Astrea-R8-Chat-9B-GGUF is a text generation model from Michionlion. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as apache-2.0.
This is an unofficial community Q80 GGUF conversion of Altworld/Astrea-R8-Chat-9B for llama.cpp, with an explicit non-reasoning chat template.
Downloads · 30 days
62
27% of all-time downloads
All-time downloads
230
Public
Repo size
9.5 GB
Likes
0
Public
Click a slice to open those files.
.gguf9.5 GB · 100%
From the Hugging Face model README
This is an unofficial community Q8_0 GGUF conversion of Altworld/Astrea-R8-Chat-9B for llama.cpp, with an explicit non-reasoning chat template.
| File | Quantization | Size |
|---|---|---|
Astrea-R8-Chat-9B-Q8_0.gguf | Q8_0 | 9,527,501,280 bytes (8.87 GiB) |
The GGUF contains a hard non-reasoning Jinja template in
tokenizer.chat_template; no external template file is required.
Astrea's optional thinking mode was not reliable in local llama.cpp testing:
simple prompts could consume hundreds of tokens before emitting </think>, and often did not end reasoning at all, and simply responded as if reasoning was not enabled.
The bundled template therefore always places a closed, empty thinking block in
the prompt and does not expose an enable_thinking template variable. It also
omits hidden reasoning when replaying assistant messages into conversation
history. A standalone copy is included as chat_template.jinja for inspection.
llama-server.exe `
--model Astrea-R8-Chat-9B-Q8_0.gguf `
--jinja `
--reasoning off `
--reasoning-format none `
--ctx-size 32768 `
--n-gpu-layers all `
--temp 0.8 `
--top-p 1.0 `
--top-k 0 `
--min-p 0.025 `
--repeat-penalty 1.08
The model metadata advertises a 262,144-token context window. Choose a context
size appropriate for your available VRAM/RAM. The command above starts at a
more conservative 32,768 tokens. I was able to easily run a much more ambitous setup with -ngl all --fit off -c 147456 -np 4 --kv-unified on a 16GB VRAM card (5070 Ti).
convert_hf_to_gguf.py using --no-mtp. The downloaded checkpoint did not
contain the extra MTP-layer tensors declared by its configuration.llama-quantize using Q8_0.gguf_new_metadata.py embedded the hard non-reasoning template;
this metadata-only copy did not requantize tensors.The tensor-only SHA-256 reported by llama-gguf-hash was identical before and
after the metadata rewrite:
20d213a0c5ee663ef6d02ffcff8d0b28cbff18b559c67aca5250cd5e6a22d624
The final whole-file checksums are in SHA256SUMS.
The final GGUF was loaded directly by llama-server without
--chat-template-file. Its exposed template matched the bundled standalone
Jinja, and a request that explicitly supplied enable_thinking=true still
returned normal content with no reasoning_content.
The source model is released under Apache-2.0. See LICENSE and NOTICE, and
refer to the source model card
for its intended use, evaluation results, and limitations.