Downloads · 30 days
0
nollafox/smilingfox
smilingfox is a image classification model from nollafox. Use it when you need a label for an image. It is set up for timm. The card lists the license as mit.
<p align="center" <img src="https://github.com/nollafox/smilingfox/raw/main/docs/banner.png" alt="smilingfox banner" style="border-radius: 16px; box-shadow: 0 8px 32px rgba(0, 0, 0, 0.12); max-width: 100%; height: aut…
Downloads · 30 days
0
Access
Public
Updated Jul 11, 2026
Repo size
483 MB
Likes
0
Public
Click a slice to open those files.
.pt483 MB · 99%
From the Hugging Face model README
This repository contains a Hugging Face–ready model bundle for an E621-focused image tagger based on SmilingFox. The bundle includes the trained weights, calibration thresholds, label inventory, and a small example script for running inference over a directory of images.
This model predicts multi-label tags for images in the E621 domain. It is designed for image tagging workflows where you want a compact, ready-to-run tagger that can be downloaded from Hugging Face Hub or used locally from this repository. Instead of a single global confidence cutoff, each of its 32,617 labels has its own calibrated decision threshold.
The repository root contains the model bundle assets needed for inference:
best.pt — model weights (fine-tuned from SmilingWolf/wd-swinv2-tagger-v3)best_thresholds.json — label list, per-label thresholds, calibration score, and decoder settingslabels.json / label_norms.json — the 32,617-label vocabulary (identical raw and normalized label lists)args.json — training and calibration configuration used to produce this checkpointconfig.json — Hugging Face transformers-style model configexample.py — inference script that tags a folder of imagesrequirements.txt — Python dependencies for example.pyFrom args.json and config.json:
SmilingWolf/wd-swinv2-tagger-v3 (via timm)asl_gamma_pos=0.0, asl_gamma_neg=4.0, asl_clip=0.05)official_ontology_loss = [0.02, 0.04, 0.06]best_thresholds.json is the source of truth for how this checkpoint was picked and how tags get decided at inference time. It holds one aggregate score plus per-label thresholds, not a micro-F1/macro-F1/precision@k/recall@k breakdown — so what follows is what's actually verifiable from the bundle, not a full benchmark report.
The best checkpoint is epoch 7 of the 10 trained, chosen on val_decoder_score = 0.2571. That score is a decoder-level objective: predictions get decoded per-image using the per-label thresholds, capped at 128 tags/image, then expanded through the label ontology's implication graph before scoring. It isn't micro-F1 or macro-F1, so don't quote it as one.
The calibration policy (support_shrunk_per_label_f1_plus_topk_then_implication_closure) picks each label's threshold by grid search — 37 steps from 0.05 to 0.95 — to maximize that label's own F1, then shrinks the result back toward the 0.35 default based on how much support the label has (shrinkage_strength=64, min_positives_for_per_label=3), so rare labels don't end up with an overconfident, narrow threshold. At decode time, predictions are capped at 128 tags/image and then closed over the ontology (predicting a child tag implies its parents).
Across all 32,617 labels, the thresholds land like this:
| Stat | Value |
|---|---|
| Min | 0.241 |
| p5 | 0.341 |
| p25 / median / p75 | 0.350 (= default) |
| p95 | 0.436 |
| Max | 0.769 |
| Mean | 0.361 |
| Std dev | 0.041 |
About 69.5% of labels (22,657 of 32,617) sit exactly on the 0.35 default — these are the lower-support labels where shrinkage pulled the calibrated value all the way back. The rest diverge meaningfully, from 0.241 for well-supported, easy-to-separate labels up to 0.769 for labels that needed a high bar to keep false positives down.
Micro-F1, macro-F1, precision@k, and recall@k aren't stored anywhere in this bundle. Getting those numbers means running example.py (or loading best.pt directly) against a labeled validation set and computing them yourself, using the per-label thresholds here as the decision boundaries.
The easiest way to use this model is through smilingfox, a Python package that wraps this repository with a proper API and CLI:
pipx install smilingfox
smilingfox tag ./photos --stdout
from smilingfox import Tagger
tagger = Tagger.from_pretrained() # downloads and caches this repo
tagger.predict("photo.jpg").tags
Tagger.from_pretrained() downloads and caches this exact repo by default, so no extra configuration is needed. See the smilingfox repository for the full API and CLI reference, including batching, custom confidence floors, and local bundle overrides.
If you'd rather not install a package, this repository's example.py works standalone against the files checked in here.
Install dependencies:
pip install -r requirements.txt
Run the example on a folder of images:
python example.py /path/to/images
Optional: download the bundle from a Hugging Face repo when it is not already present locally:
python example.py /path/to/images --hf-model-id nollafox/smilingfox
A tag is included in the output only if its score is at or above max(calibrated_tag_threshold, --required-confidence).
The script writes a JSON sidecar next to each processed image. Each sidecar contains the selected tags, for example:
{
"tags": ["tag_a", "tag_b"]
}