Downloads · 30 days
74
8% of all-time downloads
resoa/garment-attributes
garment-attributes is a image classification model from resoa. Use it when you need a label for an image. It is set up for transformers. The card lists the license as apache-2.0.
Multi-label classification of fine-grained garment construction attributes from a garment crop: silhouette, length, neckline/collar/sleeve/pocket type, opening/closure type, waistline, textile pattern, finishing techn…
Downloads · 30 days
74
8% of all-time downloads
All-time downloads
935
Public
Parameters
93.1M
372 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors372 MB · 100%
From the Hugging Face model README
Multi-label classification of fine-grained garment construction attributes from a garment crop: silhouette, length, neckline/collar/sleeve/pocket type, opening/closure type, waistline, textile pattern, finishing techniques, and fabric appearance. SigLIP2 vision encoder (Apache 2.0) with a classification head, fine-tuned on per-instance attribute annotations from Fashionpedia (CC BY 4.0).
Label space: the 218 Fashionpedia attributes with ≥100
training instances (the full list with supergroups and support counts ships
in label_space.json and in config.json id2label). Attributes map to
tech pack fields — e.g. opening type → construction/closure, textile
pattern → BOM/fabric — per the pipeline's docs/taxonomy.md.
Loads with stock transformers — no remote code:
from transformers import AutoImageProcessor, AutoModelForImageClassification
import torch
from PIL import Image
model = AutoModelForImageClassification.from_pretrained("resoa/garment-attributes")
proc = AutoImageProcessor.from_pretrained("resoa/garment-attributes")
img = Image.open("garment_crop.jpg") # crop of ONE garment, not a full scene
probs = torch.sigmoid(model(**proc(images=img, return_tensors="pt")).logits)[0]
for i, p in enumerate(probs):
if p > 0.5:
print(model.config.id2label[i], round(p.item(), 3))
garment-detector-seg).Important: input must be a single-garment crop. Full-scene inputs degrade accuracy sharply; that is what the detector stage is for.
Out of scope: fiber content, GSM, measurements, stitch class, color (color is measured deterministically in the pipeline, not classified).
google/siglip2-base-patch16-224 (Apache 2.0).scripts/prepare_attribute_data.py; 8% bbox padding, min crop 48 px,
attributes with 100+ train instances).problem_type="multi_label_classification").scripts/train_attributes.py with configs/attributes_mac.yaml
(4 epochs, batch 32, lr 2e-5, cosine schedule, trained on Apple-silicon
MPS in ~4.5 hours; val metrics improved monotonically each epoch).| Metric | Value |
|---|---|
| macro mAP | 0.442 |
| micro F1 @ 0.5 | 0.710 |
| macro F1 @ 0.5 | 0.306 |
| eval instances | 3,711 |
| evaluable labels | 212 of 218 |
Read the per-attribute table (per_attribute.md, shipped in this repo)
before using any single attribute for QC decisions: performance varies
widely by attribute, and rare attributes (support near the cutoff) can be
substantially weaker. We publish the full table, including the bad rows,
on purpose.
label_space.json is the bridge.Weights: Apache 2.0. Data: Fashionpedia, CC BY 4.0 — cite Jia et al., ECCV 2020. Base model: SigLIP2 (Google, Apache 2.0), Tschannen et al., 2025.
The warning above ("full-scene inputs degrade accuracy sharply") has now been quantified on this model's own evaluation protocol — Fashionpedia val2020, 3,711 instances, 212 evaluable labels, 48-pixel minimum crop. Only the crop padding varies; preprocessing, threshold and metric are unchanged.
| padding | micro-F1 | macro-mAP | mAP retained |
|---|---|---|---|
| 0.08 (as trained) | 0.6995 | 0.4229 | 100% |
| 0.25 | 0.7006 | 0.4355 | 103% |
| 0.50 | 0.6522 | 0.3842 | 91% |
| 1.00 | 0.5441 | 0.2403 | 57% |
| 2.00 | 0.4307 | 0.1260 | 30% |
| full scene | 0.3550 | 0.0583 | 14% |
"Sharply" is −86%: a full scene retains 14% of trained-crop mAP.
1. Widen the crop to 0.25 padding. It beats the 0.08 used in training by +3.0% macro-mAP (+0.0132 absolute; bootstrap over 3,711 instances, 400 resamples, P(>0) = 0.98).
Where that gain comes from, per instance category (same bootstrap; only categories with n >= 100 shown, and only three of thirteen show a real effect):
| category | n | Δ mAP at 0.25 | P(>0) | |
|---|---|---|---|---|
| collar | 106 | +0.0644 | 0.98 | real |
| sleeve | 1021 | +0.0301 | 0.99 | real |
| skirt | 160 | +0.0281 | 0.98 | real |
| dress, coat, neckline, lapel, pocket, pants, jacket, shorts, top | — | ≈0 | 0.19–0.83 | not distinguishable from noise |
So 0.25 is the right global default, but the benefit is concentrated in sleeve, collar and skirt crops and is neutral elsewhere — it is not worth tuning padding per category.
2. Average over several crops. Running the model on more than one padding of the same instance and combining:
| combination | micro-F1 | macro-mAP | vs 0.08 |
|---|---|---|---|
| 0.08 single | 0.6995 | 0.4229 | — |
| 0.08 + 0.25 (elementwise max) | 0.7181 | 0.4438 | +4.9% |
| 0.08 + 0.25 + 0.50 (mean) | 0.6990 | 0.4447 | +5.2% |
| 0.08 + 0.25 + 0.50 + 1.00 (max) | 0.6717 | 0.4128 | −2.4% |
Two crops at 2× inference cost gives most of the gain. Do not include the 1.00 padding — past the 0.50 cliff it costs more than it adds.
Measurement caveat: this harness reproduces the published protocol's instance and label counts exactly (3,711 / 212) but scores ~1–2% below the headline figures (0.6995 vs 0.7100 micro-F1), most likely crop-rounding or resize detail. The comparisons above are all relative to this harness's own 0.08 baseline, so that offset cancels; read the absolute numbers as ~2% low.
resoa/garment-crop-gate-nano
is a 47K-parameter, 188 KB pre-flight gate that predicts whether a crop is tight
enough to trust. On variably-cropped input it lifts macro-mAP from 0.3124 to
0.5012 (+60%) by declining roughly half the inputs — 96% of what a perfect gate
would deliver. It refuses; it does not repair.
Fashionpedia annotates instances, and its 46 categories include garment parts
(sleeve, pocket, collar, lapel, neckline) alongside whole garments
(dress, coat, jacket, pants). Training crops therefore come at two very
different scales. Median instance area as a fraction of the source image:
| part instances | whole-garment instances | ||
|---|---|---|---|
| 0.3% | jacket | 14.9% | |
| collar | 0.9% | dress | 21.9% |
| sleeve | 2.6% | coat | 23.3% |
| lapel | 4.1% | pants | 10.0% |
Two consequences worth knowing:
1. "Crop the garment first" is ambiguous. A welt (pocket) label was learned
from crops of pockets, roughly 0.3% of the source frame — not from whole
garments containing pockets. Feeding a whole jacket and reading the pocket-type
outputs asks the model about a scale it mostly saw in isolation.
2. The nickname attributes never co-occur. Measured over 206,212 training
annotations, instances carry a mean of 1.012 nickname attributes, and
cross-group co-occurrence is exactly zero — no instance has both a sleeve
style and a pocket style. That is not a fact about garments (a blazer has both);
it follows from parts being annotated as separate objects. Treat the nickname
group as one attribute of whichever part this crop depicts, not as a set of
independent garment properties.
A derived 13-group sub-taxonomy of the 96 nickname attributes — sleeve, pocket,
collar, lapel, dress, top, jacket, coat, pants, skirt, shorts, shirt, t-shirt —
is published in nickname_taxonomy.json. Each group maps 1:1 to a Fashionpedia
category, and within-group exclusivity runs 0.944–1.000.
resoa/garment-attributes-v2 is this model
retrained with crop-scale augmentation (25% padding, ±10% jitter, horizontal flip) — same
architecture, label space and hyperparameters otherwise.
Macro-mAP 0.4355 → 0.4850 (+11.4% relative, P(better) = 1.00), micro-F1 0.7006 → 0.7244, on this model's own val protocol. It wins at every evaluation padding and is markedly more forgiving of loose crops (0.4647 vs 0.3842 at 0.50 padding). 172 of 212 labels improve, with the gain inversely proportional to training support — +0.0834 mean AP for attributes with 100–300 training instances against +0.0197 for those above 5,000.
Ablation attributes the gain: scale jitter +0.0302, padding value +0.0191. The augmentation matters more than the padding value it was introduced alongside.
metal (5 mm) scores AP 0.072 despite 1,135 training examples, and
bead(a) 0.133 despite 2,030. No amount of data fixes a one-pixel feature.per_attribute_v2.json — tuxedo (jacket)
1.000 → 0.200, collarless 0.688 → 0.511.