Downloads · 30 days
0
priyadeepjaiswal9c/tiny-kws
tiny-kws is a audio classification model from priyadeepjaiswal9c. Use it for the audio classification task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A 119,372-parameter (~0.48 MB fp32) depthwise-separable CNN for spoken command recognition, trained from scratch in PyTorch. Input: 1-second 16 kHz audio → 64×101 log-mel spectrogram. Output: one of 12 classes — the k…
Downloads · 30 days
0
Access
Public
Updated Jun 13, 2026
Repo size
1 MB
Likes
0
Public
Click a slice to open those files.
.pt508 KB · 86%
From the Hugging Face model README
A 119,372-parameter (~0.48 MB fp32) depthwise-separable CNN for spoken command recognition, trained from scratch in PyTorch. Input: 1-second 16 kHz audio → 64×101 log-mel spectrogram. Output: one of 12 classes — the keywords yes, no, up, down, left, right, on, off, stop, go, plus unknown and silence.
| metric | value |
|---|---|
| accuracy | 96.65% |
| macro-F1 | 96.64% |
| CPU latency (batch=1, 1 thread, Apple M2) | 1.90 ms mean / 2.08 ms p95 |
Per-class F1 ranges from 0.921 ("unknown", the hardest class) to 0.998
("silence"); all 10 keywords score ≥0.94. Full per-class table and the
confusion matrix: see metrics.json and confusion_matrix.png in this repo.
Evaluating this checkpoint on the Colab T4 and on an Apple M2 produced
bit-for-bit identical metrics (reproducible across devices).
import torch
from huggingface_hub import hf_hub_download
# model.py + common.py from https://github.com/priyadeepjaiswal9c/tiny-kws
from model import DSCNN
from common import LogMel, normalize
ckpt = torch.load(hf_hub_download("priyadeepjaiswal9c/tiny-kws", "best.pt"),
map_location="cpu", weights_only=True)
model = DSCNN(**ckpt["model_config"]); model.load_state_dict(ckpt["model_state"]); model.eval()
wav = torch.zeros(16000) # your 1 s, 16 kHz, mono float32 waveform
feats = normalize(LogMel()(wav), ckpt["stats"])
probs = model(feats).softmax(1)[0]
print(dict(zip(ckpt["labels"], probs.tolist())))
Demo/educational model for isolated 1-second command words in quiet-to-mild noise. Not a streaming/wake-word system (no sliding-window detection), not robust to far-field audio or heavy noise, English only, and trained on crowdsourced speech that skews toward certain accents — expect degraded accuracy outside that distribution.
Live demo: https://huggingface.co/spaces/priyadeepjaiswal9c/tiny-kws · Code: https://github.com/priyadeepjaiswal9c/tiny-kws