Downloads · 30 days
42
0% of all-time downloads
lavaman131/cartoonify
cartoonify is a text-to-image model from lavaman131. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as creativeml-openrail-m.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
42
0% of all-time downloads
All-time downloads
9K
Public
Parameters
860M
26.7 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors5.5 GB · 73%
From the Hugging Face model README
This is a dreambooth model derived from runwayml/stable-diffusion-v1-5 with additional fine-tuning of the text encoder. The weights were trained from a popular animation studio using DreamBooth. Use the tokens disney style in your prompts for the effect.
You can find some example images below:
<p float="left"> <img width=256 height=256 src="./images/king.png"> <img width=256 height=256 src="./images/legend_of_zelda.png"> <img width=256 height=256 src="./images/pony.png"> <img width=256 height=256 src="./images/princess.png"> <img width=256 height=256 src="./images/red_ferrari.png"> </p>import torch
from diffusers import StableDiffusionPipeline
# basic usage
repo_id = "lavaman131/cartoonify"
device = torch.device("cuda")
torch_dtype = torch.float16 if device.type in ["mps", "cuda"] else torch.float32
pipeline = StableDiffusionPipeline.from_pretrained(repo_id, torch_dtype=torch_dtype).to(device)
image = pipeline("PROMPT GOES HERE").images[0]
image.save("output.png")
The full source-code used for training and local gradio demo for image to disney character style transfer can be found here.
As with any diffusion model, playing around with the prompt and classifier-free guidance parameter is required until you get the results you want. Zoomed-out subjects seem to loose clairity in the face. For additional safety in image generation, we use the Stable Diffusion safety checker.
The model was fine-tuned for 3500 steps on around 200 images of modern Disney characters, backgrounds, and animals. The ratios for each were 70%, 20%, and 10% respectively on an RTX A5000 GPU (24GB VRAM).
The training code used can be found here. The regularization images used for training can be found here.