Downloads · 30 days
29
40% of all-time downloads
Altworld/Astrea-R8-Chat-9B-FP8
Astrea-R8-Chat-9B-FP8 is a image-text-to-text model from Altworld. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This is the portable FP8 release of Altworld/Astrea-R8-Chat-9B, Altworld's open 9B creative-writing and chat model.
Downloads · 30 days
29
40% of all-time downloads
All-time downloads
72
Public
Parameters
9.4B
11.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors11.9 GB · 100%
How the weights are stored.
F8_E4M36.9B · 74%
From the Hugging Face model README
This is the portable FP8 release of Altworld/Astrea-R8-Chat-9B, Altworld's open 9B creative-writing and chat model.
Try Astrea · API · Documentation
The checkpoint uses LLM Compressor's compressed-tensors format and can be
served directly by compatible versions of vLLM. It contains six safetensor
shards totalling 11.09 GiB, compared with 17.53 GiB for the BF16 weights.
The checkpoint was generated from the released merged BF16 model with:
FP8_DYNAMIC, using E4M3 weights and BF16 per-output-channel scales;All 248 two-dimensional language-model linear weights are stored in FP8. The following sensitive or non-linear components remain in BF16:
lm_head;No calibration or training dataset was used during conversion.
Use a recent vLLM release with compressed-tensors support and an accelerator
with supported FP8 kernels:
vllm serve Altworld/Astrea-R8-Chat-9B-FP8 \
--max-model-len 32768 \
--gpu-memory-utilization 0.90
The quantization configuration is embedded in config.json; no separate
--quantization argument should be necessary.
Recommended sampling for writing:
{
"temperature": 0.8,
"min_p": 0.025,
"repetition_penalty": 1.08
}
Use temperature: 0.2 for ordinary factual chat. The bundled chat template
works without a system prompt and defaults to hidden thinking.
Before upload, this artifact passed the following local checks:
AutoConfig, AutoTokenizer, and AutoProcessor load successfully;Qwen3_5ForConditionalGeneration loads the compressed checkpoint and
completes a text-generation smoke test.The local Transformers smoke test dequantizes weights for its CPU fallback and therefore does not measure FP8 serving speed. Quantization changes model numerics, so applications should evaluate this checkpoint on their own prompts before replacing the BF16 release. Quant-specific Altworldbench results are not claimed on this card.
See the BF16 model card for training lineage, benchmark methodology, samples, usage guidance, and limitations. In particular:
Apache-2.0, matching the source model. See LICENSE and NOTICE.