Downloads · 30 days
192
86% of all-time downloads
kavi-labs/Krea-2-Turbo-Diffusers-FP8
Krea-2-Turbo-Diffusers-FP8 is a text-to-image model from kavi-labs. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as other.
Downloads · 30 days
192
86% of all-time downloads
All-time downloads
224
Public
Parameters
12.8B
17.9 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors17.9 GB · 100%
How the weights are stored.
F8_E4M312.8B · 100%
From the Hugging Face model README
This repository provides a native, drop-in Hugging Face Diffusers pipeline for krea/Krea-2-Turbo quantized to FP8 using PyTorch's official optimization library TorchAO.
All samples generated with 8 steps, guidance scale 0.0 at 1024x1024 on an NVIDIA RTX 4090:
<div align="center"> <img src="samples/krea2_fp8_grid.jpg" width="800" alt="Krea 2 Turbo TorchAO FP8 4-Grid Showcase" /><br/> <br/> </div> <details> <summary><b>🔍 Click to view prompt details for each quadrant:</b></summary> <br/>Top-Left (Iced Matcha Latte):
"A crisp, condensation-covered glass of iced matcha latte on a modern concrete kitchen counter, soft morning side-light, commercial food photography, 8k resolution."
Top-Right (Red Fox in Snow):
"A vibrant red fox sitting in crisp white snow at golden hour, photorealistic, sharp focus, 8k resolution."
Bottom-Left (Minimalist Workspace):
"A sleek silver laptop on a minimalist wooden desk with a ceramic coffee mug and a small succulent, soft diffused daylight."
Bottom-Right (Japanese Rock Garden):
"A serene Japanese rock garden surrounded by autumn red maple leaves, misty morning atmosphere, 8k resolution photorealistic."
Krea2Pipeline.from_pretrained(...) with native Diffusers.Float8DynamicActivationFloat8WeightConfig(granularity=PerRow()).Float8WeightOnlyConfig(), reducing text encoder footprint from 8.88 GB to 4.40 GB.num_inference_steps=8, guidance_scale=0.0).Ensure you have the latest PyTorch, Diffusers, Transformers, and TorchAO installed:
pip install --upgrade torch torchvision --extra-index-url https://download.pytorch.org/whl/cu124
pip install git+https://github.com/huggingface/diffusers.git
pip install "transformers>=4.49.0" torchao accelerate safetensors sentencepiece
import torch
from diffusers import Krea2Pipeline
# Load the official FP8 pipeline directly from Hugging Face
pipe = Krea2Pipeline.from_pretrained(
"kavi-labs/Krea-2-Turbo-Diffusers-FP8",
torch_dtype=torch.bfloat16
)
pipe.to("cuda")
# Optional: If running on a GPU with tight memory (<24 GB), use CPU offload instead:
# pipe.enable_model_cpu_offload()
# Optional: Further accelerate inference with torch.compile
# pipe.transformer = torch.compile(pipe.transformer, mode="max-autotune")
# Generate an image in 8 steps
prompt = (
"A crisp, condensation-covered glass of iced matcha latte on a modern "
"concrete kitchen counter, soft morning side-light, commercial food photography, 8k resolution."
)
image = pipe(
prompt=prompt,
width=1024,
height=1024,
num_inference_steps=8,
guidance_scale=0.0,
generator=torch.Generator("cuda").manual_seed(42)
).images[0]
image.save("matcha_latte.png")
print("✅ Image saved to matcha_latte.png")
| Component | Architecture | Original Precision (BF16) | Quantized Precision (FP8) | VRAM Footprint |
|---|---|---|---|---|
| Transformer | Krea2Transformer2DModel | 24.80 GB | TorchAO Dynamic Per-Row (W8A8) | 12.80 GB |
| Text Encoder | Qwen3VLModel (4B) | 8.88 GB | TorchAO Weight-Only | 4.40 GB |
| VAE | AutoencoderKLQwenImage | 0.50 GB | Native bfloat16 | 0.50 GB |
| Metric | Original BF16 Pipeline | This FP8 Pipeline |
|---|---|---|
| Total Model Weights | ~34.18 GB | 17.70 GB (-48% reduction) |
| Minimum Recommended GPU | 48 GB+ (A6000 / A100) | 24 GB (RTX 4090 / L40S) |
| Peak VRAM (1024x1024 Decode) | ~36.50 GB | ~19.70 GB |
| Safety Headroom (24 GB Card) | ❌ OOM (Exceeds 24 GB) | +3.82 GB Free Headroom |
Measured in live production tests at 1024x1024 resolution with 8 steps on an NVIDIA GeForce RTX 4090:
| Metric | Recorded Result | Description |
|---|---|---|
| Denoising Rate | 1.02 it/s (~8.0s total) | 8 iterative transformer steps across 13B parameters |
| Denoising + VAE Decode | ~20.1s (Cold) / ~9.8s (Warm) | Latent denoising through full image decode |
| Memory Footprint | 16.69 GB VRAM | 6.83 GB free headroom remaining on RTX 4090 |
| Encoding Speed | 11 ms | TurboJPEG SIMD compression to final output |
If you find this model conversion helpful and it saved you compute or GPU hours, consider supporting me! Every coffee helps keep open-source developer experiments and tools going:
<div align="center"> </div>