Downloads · 30 days
0
KatzenStuff/flux-2-klein-4b-webgpu
flux-2-klein-4b-webgpu is a text-to-image model from KatzenStuff. Use it when you need an image from a text prompt. It is set up for onnxruntime-web. The card lists the license as apache-2.0.
This is a community conversion of black-forest-labs/FLUX.2-klein-4B at revision e7b7dc27f91deacad38e78976d1f2b499d76a294. It packages the complete 1024×1024, four-step text-to-image path for direct browser inference w…
Downloads · 30 days
0
Access
Public
Updated Jul 25, 2026
Repo size
14.8 GB
Likes
0
Public
Click a slice to open those files.
Other8.8 GB · 60%
From the Hugging Face model README
This is a community conversion of black-forest-labs/FLUX.2-klein-4B
at revision e7b7dc27f91deacad38e78976d1f2b499d76a294. It packages the complete 1024×1024, four-step
text-to-image path for direct browser inference with ONNX Runtime WebGPU.
It is not a Diffusers or Transformers checkpoint.
The release contains the full Qwen text encoder, four FP16 transformer partitions, the FP16 VAE decoder, and immutable JSON manifests. The active payload is approximately 13.7 GiB across 378 files. Every binary referenced by the manifests has a SHA-256 and byte-length record.
The entry manifest is:
models/klein-4b/pipeline-1024/manifest.json
Manifest resource paths are relative to the manifest that declares them. A
consumer must resolve those paths against the manifest URL before fetching.
The companion browser app accepts an alternate entry manifest with its
manifest query parameter.
The source BF16 model was exported to FP16 ONNX with external data. Weights are
full and unquantized. To avoid FP16 saturation in ONNX Runtime WebGPU during
the second denoising step, the 25 residual Clip bounds were changed from
±65504 to ±60000. This is an activation-graph modification, not a weight
quantization or a LoRA.
The altered transformer partitions passed native ONNX Runtime validation against
the saved source fixture with a minimum cosine similarity of 0.9999966. The
complete browser pipeline was rendered on Apple WebGPU with finite latents and
finite decoded pixels. On the validation machine, a warm 1024×1024 four-step
generation took 26.41 seconds. Detailed reports are in validation/.
The upstream FLUX.2 Klein 4B source is Apache-2.0. This derivative includes the
upstream LICENSE.md and retains the upstream model identity and revision.
Black Forest Labs and FLUX are trademarks of their respective owners; this
community conversion is not affiliated with or endorsed by Black Forest Labs.