Downloads · 30 days
21
100% of all-time downloads
INCModel3/Z-Image-Turbo-MXFP8-RTN-AutoRound
Z-Image-Turbo-MXFP8-RTN-AutoRound is a text-to-image model from INCModel3. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as other.
This is a MXFP8 (8-bit micro-scaling) quantization of Tongyi-MAI/Z-Image-Turbo, a 6B S3-DiT distilled text-to-image model. Generated by AutoRound with RTN (round-to-nearest, iters=0).
Downloads · 30 days
21
100% of all-time downloads
All-time downloads
21
Public
Repo size
14.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors14.7 GB · 100%
From the Hugging Face model README
This is a MXFP8 (8-bit micro-scaling) quantization of Tongyi-MAI/Z-Image-Turbo, a 6B S3-DiT distilled text-to-image model. Generated by AutoRound with RTN (round-to-nearest, iters=0).
Evaluated with vllm-omni diffusion harness (8 steps, guidance 0.0, 1024×1024, seed 42).
| Benchmark | BF16 Baseline | MXFP8 Quantized |
|---|---|---|
| DrawBench CLIP | 31.73 | 31.79 |
| DrawBench CLIP-IQA | 70.57 | 70.41 |
| DrawBench ImageReward | 1.00 | 0.97 |
| GenEval | 0.757 | 0.760 |
MXFP8 quantization is essentially lossless vs the BF16 baseline (GenEval 0.760 vs 0.757, CLIP 31.79 vs 31.73).
from vllm_omni.entrypoints.omni import Omni
from vllm_omni.inputs.data import OmniDiffusionSamplingParams
omni = Omni(model="INCModel3/Z-Image-Turbo-MXFP8-RTN-AutoRound", mode="text-to-image")
params = OmniDiffusionSamplingParams(
height=1024, width=1024, seed=42,
guidance_scale=0.0, num_inference_steps=8, num_outputs_per_prompt=1,
)
out = omni.generate("a red bench in a park", sampling_params_list=[params])
Please follow the license of the original model Tongyi-MAI/Z-Image-Turbo.
Produced with autoquant-agent — agent-driven quantize + evaluate + self-heal.