Downloads · 30 days
0
Thomas-DT/apcen-multihead-tagger
apcen-multihead-tagger is a audio classification model from Thomas-DT. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for apcen-multihead-tagger. The card lists the license as cc-by-4.0.
Efficient multi-label audio tagging with a learned, physically-grounded front-end (ERB SuperGaussian filter bank + adaptive PCEN / APCEN) and a three-head model on a shared encoder (peak-pool / modulation-spectrum / t…
Downloads · 30 days
0
Access
Public
Updated Aug 26, 2026
Repo size
21.5 MB
Likes
0
Public
Click a slice to open those files.
.pt21.5 MB · 100%
From the Hugging Face model README
Efficient multi-label audio tagging with a learned, physically-grounded front-end (ERB SuperGaussian filter bank + adaptive PCEN / APCEN) and a three-head model on a shared encoder (peak-pool / modulation-spectrum / transformer-SED) fused by a per-class gate. 5.06M parameters, trained from scratch on FSD50K — no AudioSet pretraining.
Code: https://github.com/Thomas-Durand-Texte/apcen-multihead-tagger
pip install git+https://github.com/Thomas-Durand-Texte/apcen-multihead-tagger
from apcen_multihead_tagger.model import ApcenMultiheadTagger
# No argument → downloads this checkpoint from the Hub (cached afterwards).
model = ApcenMultiheadTagger.from_pretrained()
probs = model.predict(waveform) # [B, 200] sigmoid probabilities, mono @ 44.1 kHz
print(model.labels[:3]) # ['Accelerating_and_revving_and_vroom', 'Accordion', ...]
Or from the command line — --weights defaults to this repo:
python -m apcen_multihead_tagger.infer --input clip.wav --topk 5
| file | what it is |
|---|---|
best_model.pt | the released checkpoint — EMA weights + the 200-label vocabulary, 21.5 MB |
configs/waveform_cnn-19.yaml, configs/shared.yaml | the exact architecture config (also bundled in the package) |
The checkpoint is a plain torch.save payload: model_state_dict (EMA weights),
vocabulary (200 class names, index-aligned to the logits), and the selection metadata
(epoch, val_lwlrap, val_mAP). Optimizer/scheduler state was stripped for release.
[B, 200] logits → sigmoid → per-class probabilities.fine_tune.py; full fine-tune or --freeze-backbone feature-extraction).Not intended for speaker ID, transcription, or any safety-critical use.
| metric | value |
|---|---|
| lwlrap | 0.735 |
| mAP (macro) | 0.601 |
| params | 5.06M |
Checkpoint selection: best macro-mAP epoch (65), EMA weights. Macro mAP weights every class equally, so this selection favours the long tail; the lwlrap-best epoch (70) scores 0.7350 / 0.5992 — a 0.0002 lwlrap trade for +0.002 mAP.
This is the wc19 line. A larger successor (11.46M params, analytical SuperGaussian filter bank, transformer-only head) is still training and will be published separately — it does not supersede this checkpoint yet.
Per-head (single-head readouts on the same encoder): transformer 0.734 lwlrap / 0.598 mAP, pool 0.664 / 0.497, modulation 0.638 / 0.454. Honest note: at maturity the per-class gate largely concentrates on the transformer head — the pool/modulation heads' fusion contribution is small (~+0.003 mAP); their main value is the encoder co-supervision during training.
FSD50K (Fonseca et al., 2022) — 200 labels, ~41k clips, CC-BY 4.0. Weak (clip-level) labels, 44.1 kHz. No label corrections applied to this pretrain.
Purr and Animal can both fire) — treat the
scores as a multi-label set, not a single argmax.If APCEN is central to your use, cite the adaptive-PCEN work (arXiv:2510.18206) and FSD50K.