Downloads · 30 days
19
100% of all-time downloads
liodon-ai/zero-FP8
zero-FP8 is a text generation model from liodon-ai. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
FP8 quantization of movingcastles/zero, published by Liodon AI.
Downloads · 30 days
19
100% of all-time downloads
All-time downloads
19
Public
Parameters
8.2B
9.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.4 GB · 100%
How the weights are stored.
F8_E4M36.9B · 85%
From the Hugging Face model README
FP8 quantization of movingcastles/zero, published by Liodon AI.
Quantized with llm-compressor using the
FP8_DYNAMIC scheme: weights are cast to FP8 (E4M3) per-channel ahead of time, activations are
quantized to FP8 dynamically per-token at inference time. No calibration dataset is needed for this
scheme, so the quantized weights are numerically just a direct cast of the original — no calibration-set
bias to worry about. lm_head is left unquantized (standard practice — negligible size, disproportionate
quality impact if quantized).
Original size: 16.4 GB → Quantized: 9.4 GB.
vLLM
vllm serve liodon-ai/zero-FP8
Text Generation Inference (TGI)
docker run --gpus all -p 8080:80 ghcr.io/huggingface/text-generation-inference \
--model-id liodon-ai/zero-FP8
SGLang
python -m sglang.launch_server --model-path liodon-ai/zero-FP8
FP8 execution requires an NVIDIA GPU with compute capability ≥ 8.9 (Ada/Hopper/Blackwell — RTX 40-series, L4/L40S, H100/H200, B100/B200/GB10). On older GPUs, vLLM/TGI will dequantize to run, which loses the speed/memory benefit.
@misc{liodonai_zero_fp8,
title = {zero — FP8},
author = {{Liodon AI}},
year = {2026},
howpublished = {\url{https://huggingface.co/liodon-ai/zero-FP8}},
note = {FP8 (dynamic) quantization of movingcastles/zero}
}
Quantized by Liodon AI