Downloads · 30 days
236
30% of all-time downloads
WaveCut/Z-Image-Turbo-OrbitQuant-W2A4
Z-Image-Turbo-OrbitQuant-W2A4 is a machine learning model from WaveCut. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for diffusers. The card lists the license as apache-2.0.
This repository contains a compact OrbitQuant transformer-component artifact for the source Diffusers model listed above. It is intended to be loaded into the original pipeline, not used as a standalone Diffusers pipe…
Downloads · 30 days
236
30% of all-time downloads
All-time downloads
799
Public
Parameters
2.5B
7.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.5 GB · 100%
How the weights are stored.
U82.5B · 99%
From the Hugging Face model README
This repository contains a compact OrbitQuant transformer-component artifact for the source Diffusers model listed above. It is intended to be loaded into the original pipeline, not used as a standalone Diffusers pipeline repository.
OrbitQuant is a calibration-free post-training quantization method for image and video diffusion transformers. This artifact keeps the text encoders, VAE, embeddings, timestep MLP, and final heads in the source precision by default and replaces the transformer linear projections with OrbitQuant modules.
Install OrbitQuant and the Hugging Face runtime dependencies:
pip install "orbitquant[hf,kernels]>=0.9.0"
Download this model repository as an OrbitQuant artifact, then load the source Diffusers pipeline with the quantized component patched in:
import torch
from huggingface_hub import snapshot_download
from orbitquant import load_quantized_pipeline_from_artifact
artifact_id = "WaveCut/Z-Image-Turbo-OrbitQuant-W2A4"
artifact_dir = snapshot_download(artifact_id, repo_type="model")
pipe = load_quantized_pipeline_from_artifact(
artifact_dir,
torch_dtype=torch.bfloat16,
runtime_mode="auto_fused",
)
pipe.enable_model_cpu_offload(device="cuda")
image = pipe(
prompt="A precise product photo of a red ceramic mug on a wooden desk",
height=1024,
width=1024,
num_inference_steps=10,
guidance_scale=0.0,
).images[0]
image.save("z-image-orbitquant.png")
For a safetensors source checkpoint, OrbitQuant can row-stream the denoiser into packed weights through the normal Diffusers loader. Use sequential offload by replacing the final call with pipe.enable_sequential_cpu_offload().
import torch
import orbitquant
from diffusers import DiffusionPipeline
from orbitquant import (
OrbitQuantConfig,
build_diffusers_pipeline_quantization_config,
)
qconfig = build_diffusers_pipeline_quantization_config(
OrbitQuantConfig(target_policy="auto"),
components="transformer",
)
pipe = DiffusionPipeline.from_pretrained(
"Tongyi-MAI/Z-Image-Turbo",
quantization_config=qconfig,
torch_dtype=torch.bfloat16,
)
pipe.enable_model_cpu_offload()
runtime_mode="auto_fused" is the default optimized runtime. On CUDA, the kernels extra provides the Triton packed fallback; a locally built native CUDA package is preferred automatically when installed. On MPS, build and install the native Metal package from the OrbitQuant source tree. See the OrbitQuant runtime instructions. Use runtime_mode="dequant_bf16" only as an explicit compatibility/debug reference path.
Use these settings when comparing this artifact against the BF16 source model or the visual assets below:
| Setting | Value |
|---|---|
| Pipeline | ZImagePipeline |
| Resolution | 1024x1024 |
| Inference steps | 10 |
| Guidance scale | 0.0 |
| Output | image |
| Scope | paper image target |
The compact benchmark summary records native BF16-vs-OrbitQuant evidence for the comparison matrix below. Detailed per-sample generation records are retained outside this compact artifact.
| Evidence | Value |
|---|---|
| Comparison matrix | assets/image_generation_comparison_matrix.webp |
| Paired prompt/seed count | 10 |
| BF16 source generated samples | 10 |
| BF16 source generated frames | 0 |
| BF16 source nonempty outputs | 10 |
| OrbitQuant generated samples | 10 |
| OrbitQuant generated frames | 0 |
| OrbitQuant nonempty outputs | 10 |
The matrix uses all ten prompts from image_visual_v2 at 1024x1024. BF16 and OrbitQuant use the same seed for each row.
| # | Stress case | Exact required text |
|---|---|---|
| 1 | 01 Fine-detail astrolabe | - |
| 2 | 02 Layered character composition | - |
| 3 | 03 Exact counting and choreography | - |
| 4 | 04 Dense color and object binding | - |
| 5 | 05 Nested spatial relationships | - |
| 6 | 06 Cinematic night-market panorama | - |
| 7 | 07 Editorial Latin typography | ORBIT QUANT; DATA WITHOUT CALIBRATION |
| 8 | 08 Russian Constructivist typography | КВАНТОВАЯ ОРБИТА; МОСКВА 2049; КВАНТОВАНИЕ |
| 9 | 09 Japanese typography and mixed style | 量子の軌道; 東京の未来 |
| 10 | 10 Chinese typography, reflection, occlusion | 量子轨道; 未来之城 |
orbitquantW2A4W2: 44, W3: 110, W4: 84auto_fusedauto1e-10cudatriton_cudaz_imageint4_rtn_group64_bf16_activation64rpbh0paperlargest_power_of_two_dividing_dimlloyd_max2238326The following assets are stored in this artifact and compare the BF16 base generation against the OrbitQuant generation with the same prompt and seed.

Tongyi-MAI/Z-Image-Turbof332072aa78be7aecdf3ee76d5c247082da564a6apache-2.0model.safetensors: packed OrbitQuant/INT4 module tensors.quantization_config.json: serialized OrbitQuant runtime settings.orbitquant_manifest.json: source provenance, policies, module lists, and checksums.orbitquant_codebooks.safetensors: Lloyd-Max codebooks.orbitquant_rotations.safetensors: deterministic RPBH rotation metadata.auto_fused inference requires a packed matmul kernel and fails loudly when the required kernel is unavailable. The explicit dequant_bf16 reference mode materializes dequantized weights before BF16 matmul.