Downloads · 30 days
7
8% of all-time downloads
rafay99-epic/crisp-models
crisp-models is a audio classification model from rafay99-epic. Use it for the audio classification task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
Wren is the fast, lightweight model in Crisp's filler-detection family (named after birds — keen ears, clean song; the larger, higher-accuracy sibling is Kestrel). A wren is the tiniest bird yet punches far above its…
Downloads · 30 days
7
8% of all-time downloads
All-time downloads
89
Public
Repo size
2.4 MB
Likes
0
Public
Click a slice to open those files.
.pt545 KB · 51%
From the Hugging Face model README
Wren is the fast, lightweight model in Crisp's filler-detection family (named after birds — keen ears, clean song; the larger, higher-accuracy sibling is Kestrel). A wren is the tiniest bird yet punches far above its size — fitting for a ~75k-param CNN that hits 0.94 precision at ~600× real-time.
It detects filler words ("uh", "um") in speech for the Crisp macOS app — a fast, on-device alternative to running full speech-to-text just to find fillers.
[1, 1, 64, 25] (250 ms of 16 kHz mono audio,
standardized with fixed constants MEL_MEAN=-18.5658, MEL_STD=17.9252).filler_prob — P(this chunk is a filler), 0…1..mlmodel (download directly, no unzip).
Runs on-device — Core ML schedules it across the Neural Engine, GPU, or CPU.Feature extraction (mel + chunking + the slide/merge into time ranges) lives in
the host app; the model itself is a pure chunk → probability function.
train split.silencedetect).test split)| Threshold | Precision | Recall | F1 |
|---|---|---|---|
| 0.5 | 0.914 | 0.933 | 0.924 |
| 0.7 | 0.944 | 0.894 | 0.918 |
| 0.8 | 0.959 | 0.863 | 0.908 |
Speed: ~600× real-time on CPU. Recommended operating threshold 0.7 (Crisp favors precision — a false positive cuts real speech). Main error mode: occasional confusion of short real words; breaths/laughter/music are rarely misfired.
English only. Trained on podcast audio; very different domains (heavy accents, noisy rooms) may need a higher threshold or a fine-tune. Not a transcriber — it only flags fillers.
Everything is downloadable — pick whichever you need:
| File | What |
|---|---|
Wren.mlmodel | Core ML model (Apple Neural Engine). A single file the app downloads directly. |
Wren.pt | Raw PyTorch weights (state_dict) — for your own inference/fine-tuning. |
Wren.config.json | Audio framing + normalization + input/output names + recommended threshold. |
Versioning: the version is the repo's commit count, v0.0.N (mirrors
Crisp's 0.<commits> scheme); each release is one commit + one tag. Pin an exact
version in the URL:
https://huggingface.co/rafay99-epic/crisp-models/resolve/v0.0.5/Wren.mlmodel
main always points at the latest.
Code: GPL-3.0 (Crisp). Model weights derive from PodcastFillers (CC). Credited to Syntax Lab Technology / Abdul Rafay.