Downloads · 30 days
0
kkytai/flux-dev-handler-h200
flux-dev-handler-h200 is a machine learning model from kkytai. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
High-quality image generation using FLUX.1-dev optimized for NVIDIA H200 GPU with FP8 quantization.
Downloads · 30 days
0
Access
Public
Updated Jan 4, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py14.8 KB · 68%
From the Hugging Face model README
High-quality image generation using FLUX.1-dev optimized for NVIDIA H200 GPU with FP8 quantization.
| GPU | FP8 | Time per Image | Batch Size |
|---|---|---|---|
| H200 141GB | Yes | ~3-4 seconds | 16 |
| H100 80GB | Yes | ~4-5 seconds | 4 |
| L40S 48GB | Yes | ~5-6 seconds | 2-4 |
| A100 80GB | No | ~8-10 seconds | 4 |
| GPU | FP8 Support | Mode | Notes |
|---|---|---|---|
| H200 141GB | Yes (CC 9.0) | Full GPU | Optimal performance |
| H100 80GB | Yes (CC 9.0) | Full GPU | Full FP8 support |
| L40S 48GB | Yes (CC 8.9) | Full GPU / Offload | FP8 enabled |
| A100 80GB | No (CC 8.0) | Full GPU | BF16 only |
| A10G 24GB | No (CC 8.6) | CPU Offload | BF16 only |
{
"inputs": "A professional portrait of a business person in a modern office",
"parameters": {
"num_inference_steps": 28,
"guidance_scale": 3.5,
"width": 1344,
"height": 768
}
}
{
"inputs": [
"A sunset over mountains",
"A city skyline at night",
"A peaceful forest scene"
],
"parameters": {
"num_inference_steps": 28,
"guidance_scale": 3.5
}
}
Single:
{
"image": "data:image/png;base64,..."
}
Batch:
[
{"image": "data:image/png;base64,..."},
{"image": "data:image/png;base64,..."}
]
| Parameter | Default | Description |
|---|---|---|
num_inference_steps | 28 | Denoising steps (20-50 recommended) |
guidance_scale | 3.5 | Prompt adherence (3.0-7.0) |
width | 1344 | Image width |
height | 768 | Image height |
seed | -1 | Random seed (-1 for random) |
| Variable | Default | Description |
|---|---|---|
TORCH_COMPILE_MODE | reduce-overhead | Compile mode (reduce-overhead/default/max-autotune/false) |
FIXED_BATCH_SIZE | 16 | Batch size for CUDA graphs |
NUM_INFERENCE_STEPS | 28 | Default inference steps |
ENABLE_FP8 | auto | FP8 quantization (auto/true/false) |
auto (default): Enable FP8 for GPUs with Compute Capability >= 8.9true: Force enable FP8 (will fail on unsupported GPUs)false: Disable FP8 (use BF16 only)handler.py and requirements.txtHF_TOKEN secret for gated model access| Handler | FP8 | Steps | Batch | Speed (H200) | Use Case |
|---|---|---|---|---|---|
| flux-schnell | No | 4 | 10 | ~1.5s | Quick previews |
| flux-dev-no-pulid | No | 28 | 4 | ~8s | Standard quality |
| flux-dev-h200 | Yes | 28 | 16 | ~3-4s | H200 optimized |
| flux-dev-pulid | No | 28 | 1 | ~10s | Character consistency |
FP8 (8-bit floating point) quantization is applied to the transformer weights using torchao:
from torchao.quantization import quantize_, float8_weight_only
quantize_(pipe.transformer, float8_weight_only())
Benefits:
Requirements:
VAE slicing is enabled for memory-efficient batch processing:
pipe.enable_vae_slicing()
This processes images one at a time during VAE decoding, reducing peak memory usage.