Downloads · 30 days
0
Redstonexs/kagami-24k
kagami-24k is a image classification model from Redstonexs. Use it when you need a label for an image. It is set up for onnx. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Aug 27, 2026
Repo size
2.6 GB
Likes
5
Public
Click a slice to open those files.
.onnx1.3 GB · 50%
From the Hugging Face model README
English · 中文
An EVA02-L Danbooru tagger with a 24,000-tag general vocabulary that beats
wd-eva02-large-tagger-v3 on both ranking metrics as a single model —
measured on 11,639 posts newer than every compared model's training cutoff, with
paired-bootstrap confidence intervals.
鏡 — a mirror gives back what is in front of it.
11,639 held-out posts with id > 8,600,750, after the training cutoff of every
model compared (WD 7.22M, Camie ~8.36M, PixAI 8.60M). Scored on the
intersection of all seven vocabularies (7,779 tags; 3,101 with ≥10
positives), because macro-F1 over each model's own vocabulary rewards having a
bigger vocabulary rather than being more accurate.
| model | micro-F1 | macro-F1 | macro-AP | fine AP | vocab | coverage |
|---|---|---|---|---|---|---|
| Kagami-24k | 0.6710 | 0.5138 | 0.5086 | 0.6170 | 24,000 | 0.914 |
| danbooru-tagger-v1 (our previous) | 0.6569 | 0.4889 | 0.4672 | 0.5643 | 24,000 | 0.914 |
| wd-eva02-large-tagger-v3 | 0.6369 | 0.4728 | 0.4668 | 0.5979 | 8,106 | 0.506 |
| wd-vit-large-tagger-v3 | 0.6366 | 0.4574 | 0.4576 | 0.5960 | 8,106 | 0.506 |
| wd-swinv2-tagger-v3 | 0.6380 | 0.4522 | 0.4549 | 0.5821 | 8,106 | 0.506 |
| pixai-tagger-v0.9 | 0.6060 | 0.4481 | 0.4482 | 0.5696 | 9,740 | 0.588 |
| cl_tagger-1.02 | 0.5744 | 0.3973 | 0.3960 | 0.5329 | 25,098 | 0.869 |
| camie-tagger-v2 | 0.5848 | 0.3592 | 0.3509 | 0.4609 | 30,841 | 0.858 |
Point estimates are not enough, so the release criterion was fixed in advance: Δ > 0 and the 95% CI lower bound > 0, on both ranking metrics.
| baseline | Δ macro-AP | 95% CI | Δ fine AP | 95% CI | |
|---|---|---|---|---|---|
| wd-eva02-large-v3 | +0.0418 | [+0.0393, +0.0448] | +0.0190 | [+0.0060, +0.0288] | pass |
| danbooru-tagger-v1 | +0.0414 | [+0.0392, +0.0437] | +0.0527 | [+0.0321, +0.0652] | pass |
| wd-swinv2-v3 | +0.0537 | [+0.0507, +0.0568] | +0.0348 | [+0.0125, +0.0510] | pass |
| wd-vit-large-v3 | +0.0510 | [+0.0477, +0.0541] | +0.0209 | [−0.0001, +0.0349] | see below |
Resampling is paired over images, so both systems see identical data in each replicate and per-image difficulty cancels.
The one honest exception: against wd-vit-large-v3 the fine-AP interval
grazes zero, so that specific advantage is not statistically established even
though macro-AP against the same model is decisive. Fine AP averages just three
tag groups (38 tags), so its intervals are wide. Every other comparison clears.
The F1 columns are optimistic for every model listed. The scoring harness sweeps thresholds on this set. That is applied identically to all candidates so the comparison is fair, but the absolute F1 values are not clean numbers. The AP columns are threshold-free, which is why the release criterion is AP-based.
| model | flowers | hands | worn_small | scene | miko | framing |
|---|---|---|---|---|---|---|
| Kagami-24k | 0.6796 | 0.5369 | 0.6344 | 0.5505 | 0.6670 | 0.4645 |
| wd-eva02-large-v3 | 0.6750 | 0.5131 | 0.6057 | 0.5546 | 0.6162 | 0.4301 |
| danbooru-tagger-v1 | 0.6229 | 0.4785 | 0.5915 | 0.5207 | 0.6616 | 0.4298 |
5 of 6 against wd-eva02, losing only scene by 0.0041. The previous release lost
4 of 6 — including flowers and worn_small, the small-object groups.
The previous release fine-tuned WD's encoder on ground truth alone. It won the F1 metrics and lost fine-grained AP by 0.0336 — the encoder drifted, damaging exactly the small-detail tags it should have preserved.
A 2×2 ablation recombining the two encoders with the two heads showed the drift cannot be attributed to either component: the mixed pairings collapse (macro-AP 0.1743 and 0.0613, against 0.4230 and 0.4080 for the matched pairs). An additive model fitted to three of the four cells predicts the fourth at −0.1873, an impossible value for a bounded quantity. Encoder and head had co-adapted into a representation neither of WD's parts can read, so "freeze the early blocks" was never supported by evidence and was dropped.
What replaced it was anchoring rather than immobilising. A frozen teacher
teacher = 0.35 · previous + 0.65 · mean(wd_eva02, wd_vit_large, wd_swinv2)
was precomputed over all 677,355 corpus images on the 7,800 general tags shared with WD's vocabulary. The 0.35 weight was chosen on a held-out dev split by out-of-fold cross-fitting, never on the test set. Per-tag and frequency-banded weightings were both tried and rejected — a single global weight was within 0.0009 macro-AP of the best per-tag scheme, below the bar set for adopting extra complexity.
Training then ran with every block trainable under layer-wise LR decay 0.92 (encoder 2.5e-6 at the top decaying to 3.4e-7 at the bottom, head 7e-5), losing ASL on ground truth plus a teacher term whose weight decays 1.0 → 0.3 across the run, so the anchor is strongest early when drift happens and ground truth dominates late where the post-8.6M knowledge lives. 196,000 steps at batch 12, EMA 0.9998.
The teacher scored 0.4691 macro-AP / 0.6250 fine AP on dev. Kagami reached 0.4612 / 0.6211 there — recovering 87% of the macro-AP gap and 94% of the fine-AP gap between the previous model and its teacher, well above the 50–80% that ensemble distillation typically returns.
Weights are the raw parameters, not the EMA. Both were scored on dev and differed by +0.0001 macro-AP / −0.0002 fine AP — noise — so the simpler artifact ships.
import numpy as np, onnxruntime as ort
from PIL import Image
import csv
sess = ort.InferenceSession("model_prob.onnx", providers=["CPUExecutionProvider"])
tags = [r["name"] for r in csv.DictReader(open("selected_tags.csv"))]
def prep(path, side=448):
im = Image.open(path).convert("RGB")
w, h = im.size
s = side / max(w, h)
im = im.resize((max(1, round(w*s)), max(1, round(h*s))), Image.BICUBIC)
sq = Image.new("RGB", (side, side), (255, 255, 255))
sq.paste(im, ((side-im.size[0])//2, (side-im.size[1])//2))
x = np.asarray(sq, dtype=np.float32) / 255.
x = (x - 0.5) / 0.5
return x.transpose(2, 0, 1)[::-1][None].copy() # NCHW, BGR
p = sess.run(None, {"input": prep("image.jpg")})[0][0]
for i in np.argsort(-p)[:30]:
print(f"{tags[i]:30s} {p[i]:.3f}")
Input is NCHW BGR, mean=std=0.5, white square pad — matching WD's own preprocessing, not NHWC. SmilingWolf's ONNX models publish NHWC; if you are adapting an existing WD pipeline, transpose accordingly.
Suggested thresholds: 0.37 for macro-F1, 0.54 for micro-F1. These were swept on the test set and are guidance, not calibration.
Most of the vocabulary is unvalidated. Of 24,000 tags, 14,626 (61%) have fewer than 1,000 training examples, and an 11,639-image eval set cannot check them:
| n_train | tags | measurable (≥5 eval positives) | mean AP |
|---|---|---|---|
| 100–1,000 | 13,970 | 67 | 0.4095 |
| 1,000–10,000 | 6,507 | 2,278 | 0.3826 |
| 10,000+ | 2,867 | 2,828 | 0.4714 |
67 of 13,970 tags in the 100–1,000 band are measurable here. The tail is not shown to be good or bad — it is untested. The "24,000 tags" figure describes the output layer, not 24,000 validated tags.
Trained at 448 on images stored at 640 long side, so it has only ever seen double-resampled input. Feeding native resolution is sharper than anything it trained on. A controlled test showed this storage step costs nothing measurable (macro-AP +0.0001, fine AP +0.0010, CI [−0.0021, +0.0045]).
Character and copyright tags are not included — general tags only.
Apache-2.0, derived from wd-eva02-large-tagger-v3 (Apache-2.0). The three WD
models used as distillation teachers are all Apache-2.0. Camie and cl_tagger
appear only as benchmark comparisons; no weights from either are used here.
Training artifacts, logs and validation gates:
Redstonexs/danbooru-tagger-round3-artifacts.