Skip to content

inference-optimization

Qwen3-8B-DFlash-FP8-DYNAMIC

inference-optimization/Qwen3-8B-DFlash-FP8-DYNAMIC

Qwen3-8B-DFlash-FP8-DYNAMIC is a text generation model from inference-optimization. Use it when you need the model to write or continue text. It is set up for speculators. The card lists the license as apache-2.0.

Drift8 is a quantized DFlash drafter for the Qwen3-8B target, derived from RedHatAI/Qwen3-8B-speculator.dflash. This repository contains the drafter component; it is not a standalone chat model.

Downloads · 30 days

53

100% of all-time downloads

All-time downloads

53

Public

Parameters

2.3B

3.5 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors3.5 GB · 100%

Parameter types

How the weights are stored.

BF161.2B · 54%

Try a prompt

Base models

Task
Text Generation
Library
speculators
License
apache-2.0
Created
Sep 24, 2026
Updated
Sep 24, 2026