Downloads · 30 days
113
11% of all-time downloads
aipicasso/commonart-beta
commonart-beta is a text-to-image model from aipicasso. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as apache-2.0.
Downloads · 30 days
113
11% of all-time downloads
All-time downloads
1.1K
Public
Parameters
611M
17.1 GB on disk
Likes
17
Public
Click a slice to open those files.
.pth7.4 GB · 75%
From the Hugging Face model README

This is a text-to-image model learning from CC-BY-4.0, CC-0 or CC-0 like images.
At AI Picasso, we develop AI technology through active dialogue with creators, aiming for mutual understanding and cooperation. We strive to solve challenges faced by creators and grow together. One of these challenges is that some creators and fans want to use image generation but can't, likely due to the lack of permission to use certain images for training. To address this issue, we have developed CommonArt β. As it's still in beta, its capabilities are limited. However, its structure is expected to be the same as the final version.
pip install transformers diffusers
import torch
from diffusers import Transformer2DModel, PixArtSigmaPipeline, AutoencoderKL, DPMSolverMultistepScheduler
from transformers import AutoModelForCausalLM, AutoTokenizer
# Prompts
prompt = "カラフルなお花畑。赤、青、黄、紫、ピンクなどの色とりどりの花に溢れている。"
neg_prompt=""
# Settings
device = "cuda"
weight_dtype = torch.float32
weight_dtype_te = torch.bfloat16
generator = torch.Generator().manual_seed(44)
# Load text encoder
tokenizer = AutoTokenizer.from_pretrained("cyberagent/calm2-7b")
text_encoder = AutoModelForCausalLM.from_pretrained(
"cyberagent/calm2-7b",
torch_dtype=weight_dtype_te,
device_map=device
)
# Get text embeddings
with torch.no_grad():
pos_ids = tokenizer(
prompt, max_length=512, padding="max_length", truncation=True, return_tensors="pt",
).to(device)
pos_emb = text_encoder(pos_ids.input_ids, output_hidden_states=True, attention_mask=pos_ids.attention_mask)
pos_emb = pos_emb.hidden_states[-1]
neg_ids = tokenizer(
neg_prompt, max_length=512, padding="max_length", truncation=True, return_tensors="pt",
).to(device)
neg_emb = text_encoder(neg_ids.input_ids, output_hidden_states=True, attention_mask=neg_ids.attention_mask)
neg_emb = neg_emb.hidden_states[-1]
# Important
del text_encoder
# load models
transformer = Transformer2DModel.from_pretrained(
"aipicasso/commonart-beta",
torch_dtype=weight_dtype
)
vae = AutoencoderKL.from_pretrained("madebyollin/sdxl-vae-fp16-fix", torch_dtype=weight_dtype)
scheduler=DPMSolverMultistepScheduler()
pipe = PixArtSigmaPipeline(
vae=vae,
tokenizer=None,
text_encoder=None,
transformer=transformer,
scheduler=scheduler
)
pipe.to(device)
# Generate Image
with torch.no_grad():
image = pipe(
negative_prompt=None,
prompt_embeds=pos_emb,
negative_prompt_embeds=neg_emb,
prompt_attention_mask=pos_ids.attention_mask,
negative_prompt_attention_mask=neg_ids.attention_mask,
max_sequence_length=512,
width=512,
height=512,
num_inference_steps=20,
generator=generator,
guidance_scale=4.5).images[0]
image.save("flowers.png")
pip install transformers diffusers quanto
import torch
from diffusers import Transformer2DModel, PixArtSigmaPipeline, AutoencoderKL, DPMSolverMultistepScheduler
from transformers import AutoModelForCausalLM, AutoTokenizer, QuantoConfig
# Prompts
prompt = "カラフルなお花畑。赤、青、黄、紫、ピンクなどの色とりどりの花に溢れている。"
neg_prompt=""
# Settings
device = "cuda"
weight_dtype = torch.bfloat16
weight_dtype_te = torch.bfloat16
generator = torch.Generator().manual_seed(44)
# Load text encoder
tokenizer = AutoTokenizer.from_pretrained("cyberagent/calm2-7b")
quantization_config = QuantoConfig(weights="int8")
text_encoder = AutoModelForCausalLM.from_pretrained(
"cyberagent/calm2-7b",
quantization_config=quantization_config,
torch_dtype=weight_dtype_te,
device_map=device
)
# Get text embeddings
with torch.no_grad():
pos_ids = tokenizer(
prompt, max_length=512, padding="max_length", truncation=True, return_tensors="pt",
).to(device)
pos_emb = text_encoder(pos_ids.input_ids, output_hidden_states=True, attention_mask=pos_ids.attention_mask)
pos_emb = pos_emb.hidden_states[-1]
neg_ids = tokenizer(
neg_prompt, max_length=512, padding="max_length", truncation=True, return_tensors="pt",
).to(device)
neg_emb = text_encoder(neg_ids.input_ids, output_hidden_states=True, attention_mask=neg_ids.attention_mask)
neg_emb = neg_emb.hidden_states[-1]
# Important
del text_encoder
# load models
transformer = Transformer2DModel.from_pretrained(
"aipicasso/commonart-beta",
torch_dtype=weight_dtype
)
vae = AutoencoderKL.from_pretrained("madebyollin/sdxl-vae-fp16-fix", torch_dtype=weight_dtype)
scheduler=DPMSolverMultistepScheduler()
pipe = PixArtSigmaPipeline(
vae=vae,
tokenizer=None,
text_encoder=None,
transformer=transformer,
scheduler=scheduler
)
pipe.to(device)
# Generate Image
with torch.no_grad():
image = pipe(
negative_prompt=None,
prompt_embeds=pos_emb,
negative_prompt_embeds=neg_emb,
prompt_attention_mask=pos_ids.attention_mask,
negative_prompt_attention_mask=neg_ids.attention_mask,
max_sequence_length=512,
width=512,
height=512,
num_inference_steps=20,
generator=generator,
guidance_scale=4.5).images[0]
image.save("flowers.png")
See Yahoo Flickr Creative Commons 100M dataset for more information. The information was collected circa 2014 and known to have a bias towards internet connected Western countries. Some areas such as the global south lack representation.
We used these dataset to train the diffusion transformer:
Google Cloud (Tokyo Region).
We used NVIDIA L4x8 instance 4 nodes. (Total: L4x32)
We approciate the image providers. So, we are standing on the shoulders of giants.