Downloads · 30 days
59
39% of all-time downloads
opendiffusionai/sd15vae-texttuned
sd15vae-texttuned is a machine learning model from opendiffusionai. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for diffusers.
It has been said that the sd vae is not capable of faithful text rendering, at a fundamental level. I decided to test that and it turns out to not exactly be true.
Downloads · 30 days
59
39% of all-time downloads
All-time downloads
150
Public
Parameters
83.7M
1 GB on disk
Likes
4
Public
Click a slice to open those files.
.safetensors335 MB · 100%
From the Hugging Face model README
It has been said that the sd vae is not capable of faithful text rendering, at a fundamental level. I decided to test that and it turns out to not exactly be true.
This 4channel vae can render a 16pt font with almost perfect fidelity.
This is a TECH DEMO ONLY. I have trained it exclusively on text, so normal rendering has drastically suffered as a result.
This started with the original sd vae. I generated an initial dataset of 14pt-only text images, with https://github.com/ppbrown/ai-training/blob/main/trainer/vae/gen_glyph_dataset.py ( around 4000 of them) I then trained the vae at LR 1e-5 for 28000 steps, batchsize 1, using my standard settings of lpips/vgg weights 0.08, LAP 0.02 and "rawvgg" enabled, with my train_vae.py in that same repo
I then generated a MIXED size font dataset,ranging across 14,16,18, and 20pt total images around 8000 This time I chose batch size 2, LR 2e-5.
14pt text is of course more difficult than 16pt text. 16pt, however, renders almost indistinguishable from original. The specific font used across all images, was Libertinus Serif
I did have a fairly good result at 69000 steps for 14pt. I thought it had hit a floor, but then suddenly at 93000 steps, there was a dramatic breakthrough, at least by the numbers. (and then there was another one at 157000)
!(clean 16pt)[bookimg_16.step_093000.webp]
If you want to try it out for yourself, you can use the tools in my repo, such as
trainer/cache-utils/create_imgcache_sdvae.py --vae --previewonly \
--model opendiffusionai/sd15vae-texttrained --data_root someimgdir