Downloads · 30 days
73
1% of all-time downloads
dataautogpt3/Proteus-v0.6
Proteus-v0.6 is a text-to-image model from dataautogpt3. Use it when you need an image from a text prompt. It is set up for diffusers.
Downloads · 30 days
73
1% of all-time downloads
All-time downloads
8.3K
Public
Parameters
2.6B
13.9 GB on disk
Likes
18
Public
Click a slice to open those files.
.safetensors13.9 GB · 100%
From the Hugging Face model README
I'm excited to introduce Proteus v0.6, a complete rebuild of my AI image generation model. This is the first version of the rework, focusing entirely on enhancing photorealism. While it's not aiming to be state-of-the-art, I believe it's a good step forward in producing high-quality images. Please note that this is a preliminary version, and it's not the final, fully-featured checkpoint—more improvements and features will come in future updates.
Proteus v0.6 is a total rework from the ground up. In previous versions, combining different training methods and learning rates caused the model to become unstable during large-scale training. Learning from those experiences, I've retrained the model using only the photorealism aspects of the Proteus dataset.
For now, I'm calling this new training technique Multi-Perspective Fusion.
This approach involves:
I'm hoping this method will be interesting to data scientists exploring advanced training techniques.
Here's how you can use Proteus v0.6 with the Hugging Face 🧨 diffusers library:
import torch
from diffusers import (
StableDiffusionXLPipeline,
KDPM2AncestralDiscreteScheduler,
AutoencoderKL
)
# Load VAE component
vae = AutoencoderKL.from_pretrained(
"madebyollin/sdxl-vae-fp16-fix",
torch_dtype=torch.float16
)
# Configure the pipeline
pipe = StableDiffusionXLPipeline.from_pretrained(
"dataautogpt3/Proteus-v0.6",
vae=vae,
torch_dtype=torch.float16
)
pipe.scheduler = KDPM2AncestralDiscreteScheduler.from_config(pipe.scheduler.config)
pipe.to('cuda')
# Define prompts and generate image
prompt = "a cat wearing sunglasses on the beach"
negative_prompt = ""
image = pipe(
prompt,
negative_prompt=negative_prompt,
width=1024,
height=1024,
guidance_scale=7,
num_inference_steps=50,
).images[0]
image.save("generated_image.png")
Following the approach from the first version, I plan to gradually introduce new concepts and visual styles by adding one large training batch at a time. This incremental method aims to expand the model's capabilities while keeping it stable.
If anyone is interested, I'd be open to collaborating on papers about this work. I'm looking for a team to help me publish, but I'm new to this and would appreciate any guidance.
License Options:
Given my goal to allow personal use and commercial use up to a certain revenue threshold while requiring larger entities to contact me for a separate agreement, I'm considering the following existing licenses:
For more details, see the Polyform Small Business License.
This is a personal project developed solely by me.
Citation
If you use Proteus v0.6 in your work, please cite it as:
[Alexander Rafael Izquierdo], "Proteus v0.6: Multi-Perspective Fusion," 2024.