Downloads · 30 days
0
wangkanai/flux-dev-fp8
flux-dev-fp8 is a text-to-image model from wangkanai. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as apache-2.0.
FLUX.1-dev is a state-of-the-art text-to-image generation model optimized in FP8 precision for maximum performance and reduced VRAM requirements. This repository contains the complete model weights in FP8 format, offe…
Downloads · 30 days
0
Access
Public
Updated Oct 28, 2025
Repo size
43.9 GB
Likes
5
Public
Click a slice to open those files.
.safetensors73.5 GB · 93%
From the Hugging Face model README
FLUX.1-dev is a state-of-the-art text-to-image generation model optimized in FP8 precision for maximum performance and reduced VRAM requirements. This repository contains the complete model weights in FP8 format, offering professional-grade image generation with significantly reduced memory footprint compared to FP16 variants.
FLUX.1-dev is a 12-billion parameter rectified flow transformer model for text-to-image generation. This FP8 quantized version maintains generation quality while reducing VRAM requirements by approximately 50% compared to FP16, making it accessible on consumer-grade GPUs while preserving the model's creative and prompt-following capabilities.
Key Features:
flux-dev-fp8/
├── checkpoints/
│ └── flux/
│ └── flux1-dev-fp8.safetensors # 17GB - Complete checkpoint
├── diffusion_models/
│ └── flux1-dev-fp8.safetensors # 12GB - Core diffusion model
├── text_encoders/
│ ├── t5xxl-fp8.safetensors # 4.6GB - T5-XXL text encoder (FP8)
│ ├── clip-g.safetensors # 1.3GB - CLIP-G text encoder
│ ├── clip-vit-large.safetensors # 1.6GB - CLIP ViT-Large
│ └── clip-l.safetensors # 235MB - CLIP-L encoder
├── clip/
│ └── t5xxl-fp8.safetensors # 4.6GB - T5 encoder (alternate path)
├── clip_vision/
│ └── clip-vision-h.safetensors # 1.2GB - CLIP vision model
└── README.md
Total Size: ~46GB
checkpoints/flux/): Full model with all components for direct loadingdiffusion_models/): Core image generation transformertext_encoders/): Dual encoding system for text understanding
import torch
from diffusers import FluxPipeline
# Load the FP8 model (adjust paths to your local installation)
pipe = FluxPipeline.from_single_file(
"E:/huggingface/flux-dev-fp8/checkpoints/flux/flux1-dev-fp8.safetensors",
torch_dtype=torch.float16 # Use FP16 for computation
)
# Enable memory optimizations
pipe.enable_model_cpu_offload()
pipe.enable_vae_slicing()
# Generate an image
prompt = "A serene mountain landscape at sunset, photorealistic, 8k quality"
image = pipe(
prompt=prompt,
height=1024,
width=1024,
num_inference_steps=28,
guidance_scale=3.5
).images[0]
image.save("output.png")
import torch
from diffusers import FluxPipeline
from transformers import T5EncoderModel, CLIPTextModel
# Load components separately for fine-grained control
text_encoder = T5EncoderModel.from_single_file(
"E:/huggingface/flux-dev-fp8/text_encoders/t5xxl-fp8.safetensors",
torch_dtype=torch.float8_e4m3fn
)
text_encoder_2 = CLIPTextModel.from_single_file(
"E:/huggingface/flux-dev-fp8/text_encoders/clip-g.safetensors",
torch_dtype=torch.float16
)
# Load the main diffusion model
pipe = FluxPipeline.from_single_file(
"E:/huggingface/flux-dev-fp8/diffusion_models/flux1-dev-fp8.safetensors",
text_encoder=text_encoder,
text_encoder_2=text_encoder_2,
torch_dtype=torch.float16
)
pipe.to("cuda")
# Add model paths in ComfyUI:
# Settings > System Paths > Checkpoints:
# E:\huggingface\flux-dev-fp8\checkpoints\flux
#
# Settings > System Paths > CLIP:
# E:\huggingface\flux-dev-fp8\text_encoders
#
# Load workflow:
# - Add "Load Checkpoint" node
# - Select: flux1-dev-fp8.safetensors
# - Connect to KSampler with recommended settings:
# - Steps: 20-28
# - CFG: 3.5
# - Sampler: euler
# - Scheduler: simple
# 1. Enable CPU offloading (reduces VRAM to ~16GB)
pipe.enable_model_cpu_offload()
# 2. Enable VAE slicing (for high resolutions)
pipe.enable_vae_slicing()
pipe.enable_vae_tiling() # For resolutions > 2048px
# 3. Use attention slicing (reduces memory further)
pipe.enable_attention_slicing(slice_size="auto")
# 4. Use torch.compile for speed (PyTorch 2.0+)
pipe.unet = torch.compile(pipe.unet, mode="reduce-overhead", fullgraph=True)
# Recommended generation parameters
image = pipe(
prompt=your_prompt,
height=1024,
width=1024,
num_inference_steps=28, # 20-28 recommended for quality
guidance_scale=3.5, # 3.0-4.0 optimal range for FLUX
generator=torch.manual_seed(42) # For reproducibility
).images[0]
# Generate multiple images efficiently
prompts = ["prompt 1", "prompt 2", "prompt 3"]
images = pipe(
prompt=prompts,
height=1024,
width=1024,
num_inference_steps=28,
guidance_scale=3.5
).images # Returns list of images
This FP8 version uses Float8 E4M3 quantization:
| Metric | FP16 | FP8 (This Model) |
|---|---|---|
| VRAM | ~72GB | ~24GB (active), ~16GB (offloaded) |
| Speed | Baseline | 1.5-2x faster (on supported GPUs) |
| Quality | Reference | 95-98% equivalent |
| Generation | Professional | Professional |
Apache License 2.0
This model is released under the Apache 2.0 license, allowing commercial and non-commercial use with attribution. See the LICENSE file for full terms.
If you use FLUX.1-dev in your research or projects, please cite:
@misc{flux1dev2024,
title={FLUX.1: State-of-the-Art Image Generation},
author={Black Forest Labs},
year={2024},
url={https://blackforestlabs.ai/flux-1-dev/}
}
Out of Memory Errors:
# Solution: Enable all memory optimizations
pipe.enable_model_cpu_offload()
pipe.enable_vae_slicing()
pipe.enable_attention_slicing(slice_size="auto")
Slow Generation:
# Solution: Use torch.compile (requires PyTorch 2.0+)
pipe.unet = torch.compile(pipe.unet, mode="reduce-overhead")
Quality Issues with FP8:
# Solution: Use FP16 computation with FP8 weights
pipe = FluxPipeline.from_single_file(
model_path,
torch_dtype=torch.float16 # Compute in FP16, weights stay FP8
)
Model developed by: Black Forest Labs Quantization: Community contribution Repository maintained by: Local model collection Last updated: 2025-01-28