Downloads · 30 days
33
30% of all-time downloads
laion/vocalburst-classifier-single
vocalburst-classifier-single is a audio classification model from laion. Use it for the audio classification task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
A sharper classifier specialised for clips that contain exactly one isolated vocal burst (no speech). Same taxonomy (82 classes + noburst). Softmax (single most-likely type). Use it as the "name the burst" stage after…
Downloads · 30 days
33
30% of all-time downloads
All-time downloads
109
Public
Repo size
23.8 MB
Likes
0
Public
Click a slice to open those files.
.pt11.9 MB · 100%
From the Hugging Face model README
A sharper classifier specialised for clips that contain exactly one isolated vocal burst (no speech).
Same taxonomy (82 classes + no_burst). Softmax (single most-likely type). Use it as the "name the
burst" stage after a detector has isolated a segment.
⚠️ Taxonomy patch (label mapping only, weights unchanged): the three classes originally trained at indices 79
Hand Scratching Head, 80Hand Slapsand 81Slap Faceare now labelledNo burstbecause of excessive false positives in production. All 83 logits and indices are intact. See Taxonomy patch before using raw logits.
laion/voiceclap-commercial (768-d, frozen) → MLP 1024-wide × 3 deep → 83-way softmax. ~3M params.
Trained ONLY on clean single-burst clips: Gemini-confirmed DACVAE positives (82 classes, exactly one
label, no speech) + pure Emilia no_burst clips (~15.9k). Cleaner labels, no speech confusion → sharper.
50 epochs, AdamW, cross-entropy, best-mAP checkpoint.
v2 (current): retrained with 9,000 multilingual burst-free negatives (FLEURS: Chinese, Hindi, Bengali, Arabic, Persian, Urdu, Tamil, Telugu, Vietnamese, Thai, Indonesian, Japanese, Korean, Swahili, Yoruba, Zulu, Turkish, Russian) — fixes hallucinated bursts on non-European speech.
| metric | fine (82) | coarse (16 families) |
|---|---|---|
| macro mAP | 0.68 | 0.87 |
| top-1 exact | 60% | – |
| a true label in top-3 | 88% | – |
| On isolated bursts it names the coarse family very reliably (~0.87) and the fine subtype well. |
Two stages: a frozen laion/voiceclap-commercial audio encoder turns a clip into a 768-d embedding
(encode_waveform, auto-downloaded — the repo needs no extra setup), then this small trained MLP head
maps it to 83 outputs = 82 VocalBurst classes + no_burst (taxonomy: LAION-AI/voice-taxonomies · vocalburst).
A no-burst gate: if P(no_burst) ≥ 0.5 the clip is declared burst-free (no false alarm); otherwise the top classes are returned. Clips are truncated to the first 30 s (the encoder's window).
from inference import VocalBurstClassifier
clf = VocalBurstClassifier("laion/vocalburst-classifier-single") # HF repo id, or a local checkout dir
print(clf.predict("clip.wav")) # -> {no_burst, p_no_burst, top1, predictions:[(class,prob)], group}
No burstThis section supersedes the earlier "Known caveat:
Slap Facefalse positives" note. That note recommended skippingSlap Faceas top-1 and falling through to the runner-up. The problem turned out to be broader than one class and is now fixed in the label mapping itself, so the manual skip rule is no longer needed — do not apply it on top of this.
The model was originally trained with three fine classes in the
hand_and_body_sounds family:
| index | original class name (as trained) | original group |
|---|---|---|
| 79 | Hand Scratching Head | hand_and_body_sounds |
| 80 | Hand Slaps | hand_and_body_sounds |
| 81 | Slap Face | hand_and_body_sounds |
In large-scale production use — re-classifying ~1.5 M CrisperWhisper-detected burst
timestamps from in-the-wild expressive speech — all three fired as top-1 far more often
than is plausible, and manual spot-checks showed the great majority were false
positives: the true event was usually a laugh, gasp, sigh, grunt, or simply no burst at
all. Slap Face was the worst offender, but Hand Slaps and Hand Scratching Head
behaved the same way. These three are percussive, broadband, very short events, and the
frozen embedder appears to map a lot of unrelated transient energy (mouth clicks, mic
bumps, plosives, clipping) into that corner of the space.
Accordingly, in classes.json all three labels have been replaced with "No burst",
and class_to_group.json maps "No burst" to the group no_burst.
The weights were not retrained and not altered. model.pt is bit-identical to the
previous revision. The head still emits 83 logits and every index keeps its original
position, so existing checkpoints, cached embeddings, cached logits and stored
predicted_index values all stay valid. This is purely a relabelling of the output
mapping.
argmax still returns an integer in [0, 82]. Indices 79, 80 and 81 now decode to
the string "No burst" instead of the three hand-sound names."No burst" therefore appears three times in classes.json. Do not use
classes.index("No burst") to recover an index — it will always give 79. Always keep
the integer index as the identity and use the list only for display.no_burst class, spelled
no_burst (lowercase, underscore) to keep the pre-existing gate code working:
P(no_burst) ≥ 0.5 still refers to index 82 alone.{79, 80, 81, 82}. If you want a single boolean, test membership in that set
(or test class_to_group[label] == "no_burst"), not equality against one index.classes.json.hand_and_body_sounds now contains only Finger Snaps (index 78).probs = model(emb).softmax(-1)
idx = int(probs.argmax(-1)) # keep the INDEX as the identity
label = classes[idx] # 79/80/81 -> "No burst"
is_burst = idx not in (79, 80, 81, 82)
This is an observation from downstream use on in-the-wild speech, not a formal error analysis on the model's own held-out split — the reported metrics above were computed before this relabelling and still refer to the original 83-way taxonomy.
model.pt (MLP weights, unchanged) · config.json (arch) · classes.json (83 labels, index order;
79/80/81 relabelled No burst) · class_to_group.json (fine→coarse families) · inference.py ·
example.py · requirements.txt.
Training data + embeddings: laion/vocalburst-classification. Embedder: laion/voiceclap-commercial. License: CC-BY-4.0.