Downloads · 30 days
53
100% of all-time downloads
inference-optimization/Qwen3-8B-DFlash-FP8-DYNAMIC
Qwen3-8B-DFlash-FP8-DYNAMIC is a text generation model from inference-optimization. Use it when you need the model to write or continue text. It is set up for speculators. The card lists the license as apache-2.0.
Drift8 is a quantized DFlash drafter for the Qwen3-8B target, derived from RedHatAI/Qwen3-8B-speculator.dflash. This repository contains the drafter component; it is not a standalone chat model.
Downloads · 30 days
53
100% of all-time downloads
All-time downloads
53
Public
Parameters
2.3B
3.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.5 GB · 100%
How the weights are stored.
BF161.2B · 54%
From the Hugging Face model README
Drift8 is a quantized DFlash drafter for the Qwen3-8B target, derived from RedHatAI/Qwen3-8B-speculator.dflash. This repository contains the drafter component; it is not a standalone chat model.
quant_run_manifest.json.Pair this drafter with the Qwen3-8B target and a DFlash-capable vLLM build:
vllm serve Qwen/Qwen3-8B \
--spec-model inference-optimization/Qwen3-8B-DFlash-Drift8-FP8-DYNAMIC \
--spec-tokens 7 \
--spec-method dflash
config.py provides the custom drafter configuration. The experiment's serving command and runtime patch are in provenance/evaluation/.
The manifests are included at the repository root. provenance/ contains the source drafter's captured train_command.txt, a quantization command explicitly marked as reconstructed, the quantizer and calibration source snapshot, the vLLM command and patch, both target and drafter checkpoint hashes, and the nine per-subset evaluation commands. The selected seed-0 checkpoint is the same checkpoint used in the 2026-09-24 all-subset evaluation.
The PerfectBlend preparation, prompts, and hidden-state cache remain local because the prepared prompts are not redistributable. The cache sample counts and content hashes are recorded in calibration_manifest.json; no prompts or hidden-state tensors are uploaded.
The source drafter lists Apache-2.0 licensing on its Hugging Face model card.