Downloads · 30 days
46
5% of all-time downloads
Sohailhosseini/OpenThinker3-7B-FP8
OpenThinker3-7B-FP8 is a text generation model from Sohailhosseini. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
open-thoughts/OpenThinker3-7B quantized to FP8 (8-bit weights).
Downloads · 30 days
46
5% of all-time downloads
All-time downloads
974
Public
Parameters
7.6B
8.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.7 GB · 100%
How the weights are stored.
F8_E4M36.5B · 86%
From the Hugging Face model README
open-thoughts/OpenThinker3-7B quantized to FP8 (8-bit weights).
Near-lossless, no calibration data, and it halves every Linear weight. The safe default when you care about quality and have Ada/Hopper or newer.
Caveat. Needs compute capability >= 8.9 (Ada/Hopper+) to run fast.
| Source | open-thoughts/OpenThinker3-7B |
| Scheme | FP8 (8-bit) |
| Format | compressed-tensors |
| Parameters | 7.6B |
| Size on disk | 8.7 GB |
| Compression | 1.75x smaller than the 15.2 GB source |
| Left unquantized | lm_head |
| Quantized on | RTX 3090 |
| Quantized by | Sohailhosseini |
vllm serve Sohailhosseini/OpenThinker3-7B-FP8 \
--max-model-len 32768
from vllm import LLM, SamplingParams
if __name__ == "__main__":
llm = LLM("Sohailhosseini/OpenThinker3-7B-FP8", max_model_len=32768)
out = llm.chat(
[{"role": "user", "content": "What is quantization? Answer in one sentence."}],
SamplingParams(temperature=0.6, max_tokens=512),
)
print(out[0].outputs[0].text)
Served under vLLM 0.27.1 (Python 3.12, torch 2.13.0+cu130, CUDA 13, NVIDIA RTX 4090) and driven through the OpenAI-compatible /v1/chat/completions endpoint - not merely loaded. Three prompts, greedy decoding, all three coherent.
Note the hardware: FP8 kernels need compute capability >= 8.9, so this was verified on Ada rather than the Ampere cards used to produce it. That is the same floor the caveat above describes.
Produced with HF-quantized. recipe.yaml in this repo is the exact modifier stack that was applied, and the scheme, ignored layers and hardware are in the table above.
Licence is inherited from the source model. Quantization does not change what you are permitted to do with the weights.