Downloads · 30 days
0
Rohanify/Anime-Elite-V2
Anime-Elite-V2 is a text-to-image model from Rohanify. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as mit.
Built for mid-tier devices. Runs on 2 GB VRAM. Trained on a single RTX 5080. No pretrained components anywhere in the stack.
Downloads · 30 days
0
Access
Public
Updated Jun 14, 2026
Repo size
1.6 GB
Likes
2
Public
Click a slice to open those files.
.pt532 MB · 100%
From the Hugging Face model README
Built for mid-tier devices. Runs on 2 GB VRAM. Trained on a single RTX 5080. No pretrained components anywhere in the stack.
![]() | ![]() | ![]() |
Recommended First Prompt: python inference.py --prompt "1girl, portrait, long hair, looking at viewer, red eyes" --seed 56 --steps 190 --guidance 2.4
V2 is the follow-up to Anime-Elite-V1. Same from-scratch philosophy (no pretrained VAE, no LoRA over base SD, no fine-tuning). What changed is reliability — V1 had high peaks but inconsistent average samples. V2 smooths that out with EMA-averaged weights, so the floor rises significantly and the seed-to-seed variance drops.
If you want to use this model: head to Files and versions and download one of the .pt checkpoints + inference.py. That's it.
The samples below are all from V2, single-checkpoint, no post-processing, no upscaler. Just the model output as it comes out at 96×96.
![]() | ![]() | ![]() |
![]() | ![]() | ![]() |
These are cherry-picked but not unreachable — the reference seeds below reproduce these compositions consistently across GPUs (with minor bf16 numerical drift on different architectures).
After a lot of seed sweeping, this is the config that gave the strongest results:
prompt: 1girl,portrait,long hair,<color> eyes,<color> hair
guidance: 1.8 – 2.5
DDIM steps: 50 – 200 (200 = sharper, 50 = faster)
seeds: 26, 56 (peak compositions on the reference checkpoint)
Higher guidance pushes tag adherence but can wash out colors. Stay in 1.8–2.5 for the cleanest results.
pip install torch diffusers pillow
python inference.py --ckpt ckpt_e040.pt --prompt "1girl,portrait,long hair,red eyes" --seed 26
That's it. The script auto-detects EMA weights inside the checkpoint and uses them. Output saves to out/.
For batch sampling or to scan many seeds, see seed_sweep.py linked in the repo (generates a labeled grid of N seeds in one pass).
This model uses Danbooru-style tags, comma-separated. Exact matches only — no CLIP, no natural language understanding.
1girl not girl. red hair not red hairs. blue eyes not blue eye.1girl,portrait,long hair) then add specifics.1girl, portrait, long hair, short hair, blue eyes, red eyes, green eyes, purple eyes, pink eyes, red hair, blue hair, brown hair, white hair, pink hair, purple hair, green hair, smile, blush, looking at viewer, floral background, choker, closed mouth, bangs.vocab after loading to see the full 512-tag list.Honest list:
1boy mostly produces feminine-coded outputs because the model rarely saw counter-examples.Matched tags line.ckpt_e040.pt, ckpt_e045.pt, ckpt_e050.pt — EMA-weight checkpoints from the final stretch of training. e040 is the recommended default. Each ~270 MB.inference.py — single-file CLI inference scriptpeak1.png ... peak6.png — direct model outputs, no upscalingdiffusers (UNet2DConditionModel, random init)Built solo across a couple of late nights. The hardest part wasn't the training — it was finding the right single change to make on top of V1 (turned out to be EMA, just EMA, nothing else). Hope the from-scratch approach is useful to anyone exploring small-diffusion territory.