Downloads · 30 days
361
38% of all-time downloads
sorryhyun/anima-tagger
anima-tagger is a image-to-text model from sorryhyun. Use it when you need a caption or text from an image. The card lists the license as mit.
Multi-label anime image tagger. Given an image it emits a booru-style caption in exactly the format the Anima diffusion model was trained on:
Downloads · 30 days
361
38% of all-time downloads
All-time downloads
954
Public
Repo size
433 MB
Likes
16
Trending 1
Click a slice to open those files.
.safetensors5.7 MB · 85%
From the Hugging Face model README
Multi-label anime image tagger. Given an image it emits a booru-style caption in exactly the format the Anima diffusion model was trained on:
rating, count, characters, copyrights, general tags
It is the tagger behind anime_tools (dataset autotagging, position-clause captions, the curation GUI) and the anima_lora training pipeline, and works standalone.
Only our half of the checkpoint — a few MB. The backbone weights are fetched from their own repo at load time (see the license note below).
| File | Role |
|---|---|
config.json | Backend descriptor: backbone repo / arch / input size |
vocab.json | The 2,532-tag Anima vocabulary with categories, emit order and 60 tag groups |
rules.yaml | Caption normalization: replacements, tag aliases, always-remove, clothing dedup |
groups.yaml | Softmax / sentinel tag groups (eye color, hair color, …) |
thresholds.safetensors | Per-tag inference thresholds |
sidecar.safetensors + sidecar.json | The linear sidecar head and its row map |
sidecar_metrics.json | Held-out metrics of the shipped head |
animetimm/caformer_b36.dbv4-full (134 M params, 384 px, 12,476
danbooru tags, timm). Supplies rating, characters and general tags.rules.yaml renames recovered on the way. Per-tag
thresholds come from the dbv4 card's best_threshold (median 0.32). Of the 350
tags dbv4 cannot express, 223 are covered by the sidecar and 127 are hard-disabled
with a never-fire threshold.Everything downstream of the score vector is shared post-processing: threshold gating,
softmax groups ("at most one" eye color / hair color — the winner must still clear its
own threshold), count-tag dedupe and character cap, a character-confidence floor with
original fallback, top-1 copyright collapse, and slot ordering with underscores
turned into spaces.
Two kinds of tag are left out of the head on purpose:
shiro (mignon)). Such a name means nothing outside the dataset it came from, and a
head trained on a few dozen positives fires it on any look-alike. The 15 rows left out
are listed in sidecar.json["dropped_artist_oc"]; franchise characters dbv4 lacks
stay.silver hair, light brown hair, black footwear, …).
Their positives look exactly like the live tag's, so a head over them measured
macro-F1 0.18. rules.yaml aliases: folds each onto its live name instead
(silver hair → grey hair).animetimm/*.dbv4-full is GPL-3.0-licensed and access-gated. Nothing in this repo
vendors, mirrors or redistributes it: the loader calls hf_hub_download under your
token, and accepting the upstream repo's terms is what grants the download. Run
hf auth login and click through once on
the backbone's page before
first use. The MIT license on this repo covers our part only — vocab, rules, groups,
thresholds and the sidecar head.
Threshold-free mAP on the in-house 791-image held-out split, intersection vocab, against the previous in-house PE-backbone tagger (2026-08-26):
| caformer_b36 dbv4 | previous in-house head | |
|---|---|---|
| mAP (all tags) | 0.633 | 0.297 |
| Characters | 0.964 | 0.619 |
| Tail tags (freq < 200) | 0.630 | 0.285 |
| Rating (4-way acc) | 0.905 | 0.833 |
The gap is the backbone — a frozen-feature linear probe against a model fine-tuned on all of danbooru — not the head.
Shipped sidecar head on the same split (2026-09-13): copyright macro-F1 0.81,
characters 0.86, renamed generals 0.42. For people count the caption
count-tag rule is authoritative (0.943) over the sidecar's own softmax (0.927),
which is exposed as people_count_scores only.
Install anime_tools — one line, no checkout
— or uv sync in an anima_lora checkout,
which depends on it. Both halves of the checkpoint (this repo and the gated backbone)
auto-download on first use.
curl -fsSL https://github.com/sorryhyun/anime_tools/releases/latest/download/install.sh | sh
from PIL import Image
from anime_tools.tagger import AnimaTagger # anima_lora.captioning.AnimaTagger is the same class
tagger = AnimaTagger() # defaults to models/captioners/anima-tagger-dbv4/
print(tagger.predict_caption(Image.open("image.png")))
# "nsfw, 1girl, blue archive, animal ears, black hair, blush, fox girl, halo, long hair, ..."
predict() returns the structured form instead — rating / rating_scores,
people_count (+ people_count_source), scores and thresholds over the whole vocab,
kept (the emitted positives) and groups ({group_name: winner_or_None}).
predict_batch / predict_caption_batch run one backbone forward over a list. The
constructor takes device / dtype (defaults: CUDA if available, bf16; float32 on CPU
and MPS) and character_floor.
From the command line:
python -m anime_tools.tagger.cli.main --mode predict --image image.png --show_scores
Batch-tagging a dataset is the autotag stage (make caption-autotag in anima_lora,
or the GUI's Autotag button). In ComfyUI, the anima_tagger nodes shipped in
anime_tools/comfyui/ give AnimaTaggerLoader → AnimaTaggerCaption → STRING.
Every step is a python -m module in anime_tools.tagger.cli; the full recipe is in
docs/anima_tagger.md.
python -m anime_tools.tagger.cli.main --mode build_vocab --min_freq 5 # vocab + split + groups
python -m anime_tools.tagger.cli.build_dbv4_ckpt # checkpoint dir + card thresholds
python -m anime_tools.tagger.cli.train_sidecar --ckpt_dir models/captioners/anima-tagger-dbv4
train_sidecar caches one backbone forward per image, trains the head (BCE + CE, best
epoch by val mAP), F1-calibrates the sidecar rows' thresholds and writes the head.
--keep_artist_oc trains the OC rows too; --categories picks which vocab categories
the head covers.
@artist tags. dbv4 has no artist category and the sidecar deliberately does
not model one, so the @artists caption slot is always empty. The 92 artist tags in
the vocab are hard-disabled.original.general → safe,
questionable → nsfw).