Downloads · 30 days
1.2K
39% of all-time downloads
Altworld/Astrea-R8-Chat-9B-GGUF
Astrea-R8-Chat-9B-GGUF is a text generation model from Altworld. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Official text-only GGUF builds of Altworld/Astrea-R8-Chat-9B, a compact conversational and creative-writing model built on Qwen3.5-9B.
Downloads · 30 days
1.2K
39% of all-time downloads
All-time downloads
3.1K
Public
Repo size
29 GB
Likes
7
Trending 2
Click a slice to open those files.
.gguf29 GB · 100%
From the Hugging Face model README
Official text-only GGUF builds of Altworld/Astrea-R8-Chat-9B, a compact conversational and creative-writing model built on Qwen3.5-9B.
Try Astrea · API · Documentation
| Quant | Size | Use case |
|---|---|---|
| Q5_K_M | 6.02 GiB | Recommended balance of quality and size |
| Q4_K_M | 5.24 GiB | Smaller and faster; passed the same chat-routing smoke tests |
| Q6_K | 6.85 GiB | Higher fidelity |
| Q8_0 | 8.87 GiB | Highest-fidelity quant in this repository |
These are single-file, text-only models. The inherited vision components are not included because Astrea's visual behavior was not part of the release evaluation.
Use a current build of llama.cpp; Qwen3.5 support is recent.
llama-cli \
-hf Altworld/Astrea-R8-Chat-9B-GGUF:Q5_K_M \
-cnv \
--reasoning off \
--temp 0.8 \
--min-p 0.025 \
--repeat-penalty 1.08
For an OpenAI-compatible local server:
llama-server \
-hf Altworld/Astrea-R8-Chat-9B-GGUF:Q5_K_M \
--reasoning off \
--temp 0.8 \
--min-p 0.025 \
--repeat-penalty 1.08
Use --temp 0.2 for factual chat. The model carries its official chat
template inside the GGUF and does not require a system prompt. Reasoning is
disabled in the examples because Astrea's release behavior is direct response
without a visible thinking block.
Each file was loaded and generated from with llama.cpp. The checks covered:
Q4_K_M, the most aggressive quant here, correctly answered Good afternoon.
as normal chat and produced a coherent story from Write a short story about a dragon who collects spoons. Q5_K_M correctly recalled a name across turns.
The GGUFs declare the base architecture's 262,144-token context. Actual usable context depends on your hardware, llama.cpp build, and runtime settings.
Converted from the official merged BF16 release with llama.cpp commit
76f46ad.
The source checkpoint contains no MTP tensors, so conversion used --no-mtp.
Quantization used llama.cpp's Q4_K_M, Q5_K_M, Q6_K, and Q8_0 presets.
Astrea was trained on top of Qwen3.5-9B using multi-round LoRA supervised fine-tuning, alternating conversational and narrative rounds and checking each round against Altworldbench. See the BF16 model card for benchmarks, methodology, examples, limitations, and the evaluation protocol.
Astrea is English-first and can still hallucinate. Verify high-stakes claims independently. Engine and quantization differences can affect output, so test your own prompts and settings.
Apache-2.0. See LICENSE and NOTICE.