Downloads · 30 days
0
AmdGoose/FLUX.2-dev-transformer-int8wo
FLUX.2-dev-transformer-int8wo is a text-to-image model from AmdGoose. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as other.
This repository provides an INT8 weight-only quantized transformer for black-forest-labs/FLUX.2-dev.
Downloads · 30 days
0
Access
Public
Updated Jan 6, 2026
Repo size
70.7 GB
Likes
0
Public
Click a slice to open those files.
.bin70.7 GB · 100%
From the Hugging Face model README
This repository provides an INT8 weight-only quantized transformer for
black-forest-labs/FLUX.2-dev.
It is designed to be:
Only attention Linear layers (Q/K/V + projections) are quantized. All other components remain in BF16.
These components are automatically loaded from the base FLUX.2 model.
Full INT8 quantization of FLUX.2 introduces visible artifacts on ROCm. Quantizing only attention layers provides:
import torch
from diffusers import Flux2Pipeline, AutoModel
BASE_MODEL = "black-forest-labs/FLUX.2-dev"
ATTN_INT8 = "AmdGoose/FLUX.2-dev-transformer-attn-int8wo"
dtype = torch.bfloat16
device = "cuda" # ROCm uses "cuda" in PyTorch
transformer = AutoModel.from_pretrained(
ATTN_INT8,
subfolder="transformer_attn_int8wo",
torch_dtype=dtype,
use_safetensors=False,
).to(device)
pipe = Flux2Pipeline.from_pretrained(
BASE_MODEL,
transformer=transformer,
torch_dtype=dtype,
)
pipe.enable_attention_slicing()
pipe.vae.enable_tiling()
pipe.enable_model_cpu_offload()
image = pipe(
prompt="A realistic starter pack figurine in a blister box, studio lighting",
num_inference_steps=28,
guidance_scale=4,
height=1024,
width=1024,
).images[0]
image.save("out.png")