Downloads · 30 days
19
6% of all-time downloads
Aquiles-ai/FLUX.2-dev
FLUX.2-dev is a image-to-image model from Aquiles-ai. Use it when you need one image transformed into another. It is set up for diffusers. The card lists the license as other.
Note: This is a repackaging of the black-forest-labs/FLUX.2-dev model. Only the flux2-dev.safetensors file located in the root directory was removed, as it contained the same model as the one defined in the transforme…
Downloads · 30 days
19
6% of all-time downloads
All-time downloads
302
Public
Parameters
32.2B
113 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors113 GB · 100%
From the Hugging Face model README
Note: This is a repackaging of the black-forest-labs/FLUX.2-dev model. Only the
flux2-dev.safetensorsfile located in the root directory was removed, as it contained the same model as the one defined in thetransformer/folder. Because of this,diffuserswas loading the transformer twice, causing out-of-memory (OOM) errors. By removing this file, the duplicate loading was avoided and memory usage during inference was reduced from approximately 178 GB to 110 GB, enabling stable execution.

FLUX.2 [dev] is a 32 billion parameter rectified flow transformer capable of generating, editing and combining images based on text instructions.
For more information, please read our blog post.
FLUX.2 [dev] more efficient.We provide a reference implementation of FLUX.2 [dev], as well as sampling code, in a dedicated github repository.
Developers and creatives looking to build on top of FLUX.2 [dev] are encouraged to use this as a starting point.
FLUX.2 [dev] is also available in both ComfyUI and Diffusers.
For local deployment on a consumer type graphics card, like an RTX 4090 or an RTX 5090, please see the diffusers docs on our GitHub page.
As an example, here's a way to load a 4-bit quantized model with a remote text-encoder on an RTX 4090:
import torch
from diffusers import Flux2Pipeline
from diffusers.utils import load_image
from huggingface_hub import get_token
import requests
import io
repo_id = "diffusers/FLUX.2-dev-bnb-4bit" #quantized text-encoder and DiT. VAE still in bf16
device = "cuda:0"
torch_dtype = torch.bfloat16
def remote_text_encoder(prompts):
response = requests.post(
"https://remote-text-encoder-flux-2.huggingface.co/predict",
json={"prompt": prompts},
headers={
"Authorization": f"Bearer {get_token()}",
"Content-Type": "application/json"
}
)
prompt_embeds = torch.load(io.BytesIO(response.content))
return prompt_embeds.to(device)
pipe = Flux2Pipeline.from_pretrained(
repo_id, text_encoder=None, torch_dtype=torch_dtype
).to(device)
prompt = "Realistic macro photograph of a hermit crab using a soda can as its shell, partially emerging from the can, captured with sharp detail and natural colors, on a sunlit beach with soft shadows and a shallow depth of field, with blurred ocean waves in the background. The can has the text `BFL Diffusers` on it and it has a color gradient that start with #FF5733 at the top and transitions to #33FF57 at the bottom."
#cat_image = load_image("https://huggingface.co/spaces/zerogpu-aoti/FLUX.1-Kontext-Dev-fp8-dynamic/resolve/main/cat.png")
image = pipe(
prompt_embeds=remote_text_encoder(prompt),
#image=[cat_image] #optional multi-image input
generator=torch.Generator(device=device).manual_seed(42),
num_inference_steps=50, #28 steps can be a good trade-off
guidance_scale=4,
).images[0]
image.save("flux2_output.png")
Using the model in BF16 (Requires an H200)
import torch
from diffusers.pipelines.flux2.pipeline_flux2 import Flux2Pipeline
from transformers import Mistral3ForConditionalGeneration
from diffusers.models.transformers.transformer_flux2 import Flux2Transformer2DModel
from diffusers.models.autoencoders.autoencoder_kl_flux2 import AutoencoderKLFlux2
MODEL_ID = "Aquiles-ai/FLUX.2-dev"
text_encoder = Mistral3ForConditionalGeneration.from_pretrained(
MODEL_ID, subfolder="text_encoder", torch_dtype=torch.bfloat16, device_map="cuda"
)
dit = Flux2Transformer2DModel.from_pretrained(
MODEL_ID, subfolder="transformer", torch_dtype=torch.bfloat16, device_map="cuda"
)
vae = AutoencoderKLFlux2.from_pretrained(
MODEL_ID,
subfolder="vae",
torch_dtype=torch.bfloat16.to("cuda")
)
pipeline = Flux2Pipeline.from_pretrained(
MODEL_ID, text_encoder=text_encoder, transformer=dit, vae=vae, dtype=torch.bfloat16
).to(device="cuda")
prompt = "Realistic macro photograph of a hermit crab using a soda can as its shell, partially emerging from the can, captured with sharp detail and natural colors, on a sunlit beach with soft shadows and a shallow depth of field, with blurred ocean waves in the background. The can has the text `BFL Diffusers` on it and it has a color gradient that start with #FF5733 at the top and transitions to #33FF57 at the bottom."
output = pipeline(
prompt=prompt,
num_inference_steps=50,
generator=torch.Generator(device="cuda").manual_seed(42),
guidance_scale=4,
).images[0]
output.save("flux2_output.png")
Using the model with a quantized text encoder: Even with the text encoder quantized, the model still has high memory requirements. While it may run on an H100, the available VRAM would be very tight, so it is recommended to use GPUs with more than 85 GB of VRAM to ensure stable execution and avoid OOM issues.
import torch
from diffusers.pipelines.flux2.pipeline_flux2 import Flux2Pipeline
from transformers import Mistral3ForConditionalGeneration
from diffusers.models.transformers.transformer_flux2 import Flux2Transformer2DModel
from diffusers.models.autoencoders.autoencoder_kl_flux2 import AutoencoderKLFlux2
MODEL_ID = "Aquiles-ai/FLUX.2-dev"
MODEL_4BIT = "diffusers/FLUX.2-dev-bnb-4bit"
text_encoder = Mistral3ForConditionalGeneration.from_pretrained(
MODEL_4BIT, subfolder="text_encoder", torch_dtype=torch.bfloat16, device_map="cuda"
)
dit = Flux2Transformer2DModel.from_pretrained(
MODEL_ID, subfolder="transformer", torch_dtype=torch.bfloat16, device_map="cuda"
)
vae = AutoencoderKLFlux2.from_pretrained(
MODEL_ID,
subfolder="vae",
torch_dtype=torch.bfloat16.to("cuda")
)
pipeline = Flux2Pipeline.from_pretrained(
MODEL_ID, text_encoder=text_encoder, transformer=dit, vae=vae, dtype=torch.bfloat16
).to(device="cuda")
prompt = "Realistic macro photograph of a hermit crab using a soda can as its shell, partially emerging from the can, captured with sharp detail and natural colors, on a sunlit beach with soft shadows and a shallow depth of field, with blurred ocean waves in the background. The can has the text `BFL Diffusers` on it and it has a color gradient that start with #FF5733 at the top and transitions to #33FF57 at the bottom."
output = pipeline(
prompt=prompt,
num_inference_steps=50,
generator=torch.Generator(device="cuda").manual_seed(42),
guidance_scale=4,
).images[0]
output.save("flux2_output.png")
Black Forest Labs is committed to the responsible development and deployment of our models. Prior to releasing the FLUX.2 family of models, we evaluated and mitigated a number of risks in our model checkpoints and hosted services, including the generation of unlawful content such as child sexual abuse material (CSAM) and nonconsensual intimate imagery (NCII). We implemented a series of pre-release mitigations to help prevent misuse by third parties, with additional post-release mitigations to help address residual risks:
This model falls under the FLUX [dev] Non-Commercial License.