Downloads · 30 days
83
1% of all-time downloads
KJML/gpt-oss-20b-FP8-Dynamic
gpt-oss-20b-FP8-Dynamic is a text generation model from KJML. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This repository provides an FP8-dynamic quantized variant of OpenAI’s gpt-oss-20b model. It is intended for users who want the reasoning capabilities of gpt-oss-20b with a smaller memory footprint and faster inference…
Downloads · 30 days
83
1% of all-time downloads
All-time downloads
15.3K
Public
Parameters
20.9B
41.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors41.2 GB · 100%
How the weights are stored.
F1620.3B · 97%
From the Hugging Face model README
KJML/gpt-oss-20b-FP8-DynamicThis repository provides an FP8-dynamic quantized variant of OpenAI’s gpt-oss-20b model.
It is intended for users who want the reasoning capabilities of gpt-oss-20b with a smaller memory footprint and faster inference on modern GPUs that support FP8 inference.
⚠️ This model is not trained or fine-tuned further; it is a post-training quantization of the original
openai/gpt-oss-20bweights.
openai/gpt-oss-20bgpt-oss-20b (long-context, Harmony-format chat)openai/gpt-oss-20b (no additional training; quantization only)The original gpt-oss-20b is an open-weight reasoning model from OpenAI, designed for agentic workflows, tool use, and configurable reasoning effort. This FP8-dynamic variant preserves those capabilities while targeting more efficient deployment.
Typical direct-use scenarios (without additional fine-tuning):
Note: The model is trained on OpenAI’s Harmony response format. For best results, use a chat template that applies the Harmony format (e.g. tokenizer.apply_chat_template in Transformers) when prompting.
The FP8-dynamic variant can be used as a drop-in replacement for openai/gpt-oss-20b in:
If you fine-tune or adapt this model further, treat it as you would the base gpt-oss-20b model, but keep in mind that quantization can slightly change numeric behavior, especially for very long generations.
The model (and this quantized variant) is not recommended for:
Users should always keep a human in the loop for sensitive or impactful applications.
This model inherits all biases, risks, and limitations of the base gpt-oss-20b model. As a large language model trained on internet-scale data, it may:
The FP8-dynamic quantization may also:
Basic usage with 🤗 Transformers:
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "KJML/gpt-oss-20b-FP8-Dynamic"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto", # Will use FP8 where supported
device_map="auto",
)
messages = [
{"role": "system", "content": "You are a helpful AI assistant."},
{"role": "user", "content": "Explain what FP8 dynamic quantization is in simple terms."},
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=256,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Make sure you are using a recent version of Transformers and a PyTorch build that supports FP8 where applicable.
No new training data is introduced in this repository.
openai/gpt-oss-20b.No additional gradient-based training was performed. The steps were:
openai/gpt-oss-20b weights.safetensors format for deployment.No extra data preprocessing was done beyond what OpenAI used for the base model.
Exact performance depends on your hardware and FP8 support, but in general:
You should benchmark on your own GPU(s) for precise numbers.
No separate benchmark suite has been run specifically for the FP8-dynamic variant at this time.
openai/gpt-oss-20b, with minor differences due to quantization.If you run your own evals (e.g. on reasoning or coding benchmarks), please feel free to share issues / PRs or discussion links so others can reference them.
No additional interpretability or probing analysis has been carried out on this quantized variant.
For deeper analysis and interpretability work, refer to:
gpt-oss-20b.This repository does not involve training a new model.
For estimates of training-time emissions, please consult the original gpt-oss model card and related publications.
Architecture: Mixture-of-Experts Transformer language model (same as gpt-oss-20b)
Objective: Next-token prediction / causal language modeling
Quantization:
The quantization is applied in a way that preserves the original architecture and I/O behavior.
Quantization was performed on a single modern GPU (exact details may vary; see repository description or commits if you need exact hardware).
If you use this model in academic or commercial work, please cite at least the original gpt-oss paper/model card from OpenAI:
BibTeX:
@misc{openai2025gptoss120bgptoss20bmodel,
title={gpt-oss-120b & gpt-oss-20b Model Card},
author={OpenAI},
year={2025},
eprint={2508.10925},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2508.10925}
}
You may also optionally reference this quantized variant as:
@misc{kjml2025gptoss20bfp8dynamic,
title={KJML/gpt-oss-20b-FP8-Dynamic: FP8-dynamic Quantized Variant of gpt-oss-20b},
author={KJML},
year={2025},
howpublished={Hugging Face model repository},
url={https://huggingface.co/KJML/gpt-oss-20b-FP8-Dynamic}
}
openai/gpt-oss-20b on Hugging Face and the official gpt-oss GitHub repository.