Downloads · 30 days
22
30% of all-time downloads
laion/vocalburst-classifier-multilabel
vocalburst-classifier-multilabel is a audio classification model from laion. Use it for the audio classification task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
A presence classifier for non-speech vocal bursts (laughs, sighs, gasps, coughs, groans, throat sounds, …) over the 82-class LAION VocalBurst taxonomy + noburst. Multi-label (independent sigmoid per class) — a clip ca…
Downloads · 30 days
22
30% of all-time downloads
All-time downloads
74
Public
Repo size
115 MB
Likes
0
Public
Click a slice to open those files.
.pt57.4 MB · 100%
From the Hugging Face model README
A presence classifier for non-speech vocal bursts (laughs, sighs, gasps, coughs, groans, throat
sounds, …) over the 82-class LAION VocalBurst taxonomy + no_burst. Multi-label (independent
sigmoid per class) — a clip can contain several bursts. Works on pure bursts and bursts embedded in speech.
laion/voiceclap-commercial (768-d, frozen) → MLP 2048-wide × 4 deep (LayerNorm+GELU+dropout) → 83 sigmoids. ~14.4M trainable params.
laion/vocalburst-locator.v2 (current): retrained with 9,000 multilingual burst-free negatives (FLEURS: Chinese, Hindi, Bengali, Arabic, Persian, Urdu, Tamil, Telugu, Vietnamese, Thai, Indonesian, Japanese, Korean, Swahili, Yoruba, Zulu, Turkish, Russian) — fixes hallucinated bursts on non-European speech.
| metric | fine (82) | coarse (16 families) |
|---|---|---|
| macro mAP | 0.39 | 0.64 |
| top-1 exact | 49% | – |
| a true label in top-3 | 77% | – |
| Best when you take the top-1 prediction (which the two-stage combo demo uses). Much stronger at the coarse "which family of sound" level than the exact fine subtype. |
Two stages: a frozen laion/voiceclap-commercial audio encoder turns a clip into a 768-d embedding
(encode_waveform, auto-downloaded — the repo needs no extra setup), then this small trained MLP head
maps it to 83 outputs = 82 VocalBurst classes + no_burst (taxonomy: LAION-AI/voice-taxonomies · vocalburst).
A no-burst gate: if P(no_burst) ≥ 0.5 the clip is declared burst-free (no false alarm); otherwise the top classes are returned. Clips are truncated to the first 30 s (the encoder's window).
from inference import VocalBurstClassifier
clf = VocalBurstClassifier("laion/vocalburst-classifier-multilabel") # HF repo id, or a local checkout dir
print(clf.predict("clip.wav")) # -> {no_burst, p_no_burst, top1, predictions:[(class,prob)], group}
model.pt (MLP weights) · config.json (arch) · classes.json (83 labels, index order) ·
class_to_group.json (fine→16 coarse families) · inference.py · example.py · requirements.txt.
Training data + embeddings: laion/vocalburst-classification. Embedder: laion/voiceclap-commercial. License: CC-BY-4.0.