Downloads · 30 days
0
DeltaAtlas/Haru_no_Sumi
Haru_no_Sumi is a unconditional image generation model from DeltaAtlas. Use it for the unconditional image generation task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
An unconditional image diffusion model trained from scratch, generating illustrations in the style of Japanese woodblock and brush-and-ink prints — figures in traditional dress, muted washes on aged paper, and calligr…
Downloads · 30 days
0
Access
Public
Updated Jul 25, 2026
Repo size
2.1 GB
Likes
1
Public
Click a slice to open those files.
.safetensors2.1 GB · 100%
From the Hugging Face model README
An unconditional image diffusion model trained from scratch, generating illustrations in the style of Japanese woodblock and brush-and-ink prints — figures in traditional dress, muted washes on aged paper, and calligraphy-like marks in the margins.
The training images had no captions, which makes this an unconditional model: it does not take text prompts. Sampling draws images from the distribution the model learned rather than from a description. Adding caption data would be the first step toward conditional generation, and that is the main thing I would change on a second attempt.

Sampling is uneven. The model reliably reproduces the palette, paper texture, and brush character of the training data, and often composes a plausible figure. Global structure is weaker: faces, hands, and subject boundaries frequently dissolve, and some samples resolve into texture without a coherent subject. This is the expected failure mode for an unconditional model trained on a small dataset — style is learned well before structure.
Trained from scratch on consumer hardware. A single full training run took roughly five days.
This was the first model I trained end to end, and most of what I took away was about data rather than architecture.
Annotation matters more than I expected. Training on uncaptioned images is what makes this model unconditional. Captions are not a nice-to-have; they determine what the model is even capable of. Doing it by hand does not scale, so an automated annotation pass (captioning model or metadata extraction) would have been the single highest-value thing I could have built before training.
Dataset cleaning shows up directly in the outputs. Inconsistencies I left in the training set surfaced as artifacts and oddities in generated images. Time spent filtering and normalizing images beforehand would have paid for itself several times over.
Long training runs change how you plan. A single run took about five days, so there was no cheap way to iterate. Small decisions made before training — resolution, dataset composition, schedule — had outsized effects, because a mistake costs days rather than minutes. Next time I would validate the pipeline on a tiny subset first and reserve the long run for a configuration I had already sanity-checked.