Downloads · 30 days
15
100% of all-time downloads
INCModel3/Z-Image-Turbo-MXFP4-RTN-AutoRound
Z-Image-Turbo-MXFP4-RTN-AutoRound is a text-to-image model from INCModel3. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as other.
This is a MXFP4 (4-bit micro-scaling) quantization of Tongyi-MAI/Z-Image-Turbo, a 6B S3-DiT distilled text-to-image model. Generated by AutoRound with RTN (round-to-nearest, iters=0).
Downloads · 30 days
15
100% of all-time downloads
All-time downloads
15
Public
Repo size
11.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors11.7 GB · 100%
From the Hugging Face model README
This is a MXFP4 (4-bit micro-scaling) quantization of Tongyi-MAI/Z-Image-Turbo, a 6B S3-DiT distilled text-to-image model. Generated by AutoRound with RTN (round-to-nearest, iters=0).
Evaluated with vllm-omni diffusion harness (8 steps, guidance 0.0, 1024×1024, seed 42).
| Benchmark | BF16 Baseline | MXFP4 Quantized |
|---|---|---|
| DrawBench CLIP | 31.73 | 31.66 |
| DrawBench CLIP-IQA | 70.57 | 68.52 |
| DrawBench ImageReward | 1.00 | 0.91 |
| GenEval | 0.757 | 0.741 |
MXFP4 quantization is nearly lossless vs the BF16 baseline (GenEval 0.741 vs 0.757, CLIP 31.66 vs 31.73).
from vllm_omni.entrypoints.omni import Omni
from vllm_omni.inputs.data import OmniDiffusionSamplingParams
omni = Omni(model="INCModel3/Z-Image-Turbo-MXFP4-RTN-AutoRound", mode="text-to-image")
params = OmniDiffusionSamplingParams(
height=1024, width=1024, seed=42,
guidance_scale=0.0, num_inference_steps=8, num_outputs_per_prompt=1,
)
out = omni.generate("a red bench in a park", sampling_params_list=[params])
Please follow the license of the original model Tongyi-MAI/Z-Image-Turbo.
Produced with autoquant-agent — agent-driven quantize + evaluate + self-heal.