Downloads · 30 days
222
4% of all-time downloads
common-canvas/CommonCanvas-S-C
CommonCanvas-S-C is a text-to-image model from common-canvas. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as cc-by-sa-4.0.
Downloads · 30 days
222
4% of all-time downloads
All-time downloads
6K
Public
Parameters
866M
71.7 GB on disk
Likes
11
Public
Click a slice to open those files.
.pt14.7 GB · 53%
From the Hugging Face model README
Version Number: 0.1
CommonCanvas is a family of latent diffusion models capable of generating images from a given text prompt. Different models in the family are different sizes, and trained on different subsets of the CommonCatalog Dataset (See Data Card), a large dataset of Creative Commons licensed images with synthetic captions produced using a pre-trained BLIP-2 captioning model. CommonCanvas-S-NC is the small (S) model based off the Stable Diffusion 2 architecture, and trained on the non-commercial (NC) subset of CommonCatalog.
The goal of this purpose is to produce a high-quality text-to-image model, but to do so using an easily accessible dataset of known provenance. The exact training recipe of the model can be found in the paper hosted at this link. https://arxiv.org/abs/2310.16825
Input: CommonCatalog Text Captions
Output: CommonCatalog Images
Architecture: Stable Diffusion 2
CommonCanvas under-performs in several categories, including faces, general photography, and paintings (see paper, Figure 8). These datasets all originated from the Conceptual Captions dataset, which relies on web-scraped data. These web-sourced captions, while abundant, may not always align with human-generated language nuances. Transitioning to synthetic captions introduces certain performance challenges, however, the drop in performance is not as dramatic as one might assume.
The model is trained on 10 year old YFCC data and may not have modern concepts or recent events in its training corpus. Performance on this model will be worse on certain proper nouns or specific celebrities, but this is a feature not a bug. The model may not generate known artwork, individual celebrities, or specific locations due to the autogenerated nature of the caption data.
Note: The non-commercial variants of this model are explicitly not intended to be use
We recommend using the MosaicML Diffusion Repo to finetune / train the model: https://github.com/mosaicml/diffusion. Example finetuning code coming soon.
Try the model demo on Hugging Face Spaces
from diffusers import StableDiffusionPipeline
pipe = StableDiffusionXLPipeline.from_pretrained(
"common-canvas/CommonCanvas-S-C",
custom_pipeline="hyoungwoncho/sd_perturbed_attention_guidance", #read more at https://huggingface.co/hyoungwoncho/sd_perturbed_attention_guidance
torch_dtype=torch.float16
).to(device)
prompt = "a cat sitting in a car seat"
image = pipe(prompt, num_inference_steps=25).images[0]
We validated the model against Stability AI’s SD2 model and compared human user study
We thank @multimodalart, @Wauplin, and @lhoestq at Hugging Face for helping us host the dataset, and model weights.
@article{gokaslan2023commoncanvas,
title={CommonCanvas: An Open Diffusion Model Trained with Creative-Commons Images},
author={Gokaslan, Aaron and Cooper, A Feder and Collins, Jasmine and Seguin, Landan and Jacobson, Austin and Patel, Mihir and Frankle, Jonathan and Stephenson, Cory and Kuleshov, Volodymyr},
journal={arXiv preprint arXiv:2310.16825},
year={2023}
}