Downloads · 30 days
70
1% of all-time downloads
doohickey/doohickey-mega
doohickey-mega is a text-to-image model from doohickey. Use it when you need an image from a text prompt. It is set up for diffusers.
Models better suited for High-Resolution Image Synthesis. The main model (doohickey/doohickey-mega) has been finetuned from runwayml/stable-diffusion-v1-5 near a resolution of 768x768 (suggested method of generating f…
Downloads · 30 days
70
1% of all-time downloads
All-time downloads
6.1K
Public
Repo size
18.3 GB
Likes
4
Public
Click a slice to open those files.
.ckpt12.8 GB · 70%
From the Hugging Face model README
Models better suited for High-Resolution Image Synthesis. The main model (doohickey/doohickey-mega) has been finetuned from runwayml/stable-diffusion-v1-5 near a resolution of 768x768 (suggested method of generating from model is with Doohickey).
Current models:
| name | description | datasets used |
|---|---|---|
| doohickey/doohickey-mega/v1-3000steps.ckpt | first try, rlly good hd, bad results w/ other aspect ratios than 1:1 trained at 704x704 | A-1k |
| doohickey/doohickey-mega/v2-3000steps.ckpt | same as last one but worse | A-1k + ~1k samples from LAION-2b-En-Aesthetic >=768x768 |
| doohickey/doohickey-mega/v3-3000.ckpt | with new CLIP model (laion/CLIP-ViT-L-14-laion2B-s32B-b82K) (CLIP model also finetuned the 3k steps), models past this point were trained with various aspect ratios from 640x640 min to 768x768 max resolution. (examples 768x640 or 704x768) | A-1k + E-10k |
| doohickey/doohickey-mega/v3-6000.ckpt | 3k steps on top of v3-3000.ckpt, better at hands! (just UNet finetune, added a RandomHorizontalFlip operation at 50%) | A-1k |
| doohickey/doohickey-mega/v3-7000.ckpt | continuation of last model, I thought Colab would crash after 3k steps but it kept going for a little while saving ckpts every 1k steps. | A-1k |
| doohickey/doohickey-mega/v3-8000.ckpt | see last description, v3-6000 + 2k steps | A-1k |
The currently loaded model for diffusers is doohickey/doohickey-mega/v3-8000.ckpt
Datasets:
| name | description |
|---|---|
| A-1K | 1k scraped images, captioned with BLIP (more refined aesthetic) |
| E-10k | 10k scraped images captioned with BLIP (less refined aesthetic) |
Limitations and Biases from Stable Diffusion also apply to this model.
<div style="font-size:10px"> This model is open access and available to all, with a CreativeML OpenRAIL-M license further specifying rights and usage. The CreativeML OpenRAIL License specifies: