Downloads · 30 days
0
IOTEverythin/phoneme-wake-word
phoneme-wake-word is a automatic speech recognition model from IOTEverythin. Use it when you need speech turned into text. It is set up for pytorch. The card lists the license as apache-2.0.
A wake word here is not a model. It is ~50 bytes of configuration.
Downloads · 30 days
0
Access
Public
Updated Aug 5, 2026
Repo size
44.7 MB
Likes
3
Public
Click a slice to open those files.
.pt44.7 MB · 96%
From the Hugging Face model README
A wake word here is not a model. It is ~50 bytes of configuration.
One small streaming phoneme recognizer is trained once on generic speech and never sees a wake word. Any phrase — typed, invented, whatever you like — becomes a detector in milliseconds via grapheme-to-phoneme lookup, optionally sharpened by five spoken examples. The whole runtime fits on an ESP32-S3 at 340 KB INT8 and 14 ms per 40 ms of audio, leaving wake words changeable without retraining, reflashing, or a cloud round-trip.
| Every per-word-trained system | This | |
|---|---|---|
| New wake word costs | TTS synthesis + GPU training (minutes–hours) | a dictionary lookup (ms) |
| Artifact per word | a 50–200 KB model | ~50 bytes (phone ids + threshold) |
| N simultaneous words | N models | 1 shared model + N tiny decoders |
| Changing a word on device | reflash | send a few bytes |
Runtime lane (wake-word-agnostic, always on):
mic → 40-band log-mel + causal EMA normalization → streaming causal TCN → CTC phoneme posteriors every 20 ms → keyword-filler Viterbi decoder
Enrollment lane (per word, no training):
text → G2P → phone sequence and/or 5 spoken examples → phoneme decode → pronunciation variants → automatic per-variant threshold calibration
→ config, or an explicit refusal if the phrase cannot be separated
from ordinary speech.
The decoder is a filler-normalized Viterbi search with four gates, each closing a failure mode we hit in practice:
Detections use only within-frame logit differences, so the deployed engine never computes a softmax.
| File | Params | dev-clean PER | Common Voice dev PER | Use |
|---|---|---|---|---|
phoneme_tcn_student.pt | 358,504 | 0.229 | 0.402 | deploy this |
phoneme_tcn_teacher.pt | 3,337,768 | 0.150 | — | distillation teacher |
Student architecture — causal TCN: stride-2 stem (k=5), 8 depthwise-separable causal blocks (k=5, dilations 1,2,4,8 ×2, 192 ch, BatchNorm+ReLU, residual), 1×1 head over 40 classes (39 stress-free ARPAbet phones + CTC blank). 20 ms output frames, ~2.4 s receptive field. Every op is BatchNorm-foldable conv / ReLU / residual — chosen so post-training INT8 survives intact (99.3% frame-argmax agreement with float).
Training — CTC over phonemized transcripts (CMUdict + neural G2P
fallback) on LibriSpeech 960 h + 1.10 M Common Voice 17 English clips
(~1,500 h, incl. 83k Indian-accented). On-GPU augmentation: synthetic
room impulse responses (T60 0.1–0.6 s), additive noise (5–30 dB SNR),
same-batch babble (10–25 dB), random gain, SpecAugment. The student was
then distilled from the teacher with speech-weighted KD — frames are
weighted by 1 − p(blank) because ~75% of CTC frames are blank and
uniform KD otherwise spends its budget teaching silence (uniform KD
degraded the student; the weighted version improved it).
Ten phrases, Piper LibriTTS-R positives verified by Whisper (raw TTS is unreliable: "tornado" → "Pornado", "dakota" → "Decoder"), against 1.22 h LibriSpeech dev-clean + 0.82 h Common Voice dev negatives. Noisy conditions degrade positives and negatives identically so operating points stay condition-matched.
| Condition | Recall | Notes |
|---|---|---|
| Text-only, speaker-independent | median 0.38 @ 0 FA | hardest case: arbitrary phrase, arbitrary speaker, zero examples |
| Cross-speaker enrollment | repairs dictionary mismatch | tornado 0 → 0.35, "hey jarvis" 0.45 → 0.75 |
| Personal enrollment (5 examples) | mean 0.554, median 0.667 @ 3.3 FA/h | synthetic renditions vary more than a self-consistent human |
| Universal auto-calibration | 83% of arbitrary voices (neptuno) | 36 TTS voices, population-voted variants |
| On-device, single user | ~14/15 utterances | ESP32-S3 + INMP441, live session |
Safety property: calibration refuses phrases it cannot separate.
"norman" (one phone from "normal") was rejected for 5/6 voices and
"dakota" for 6/6 — independently flagged by the text-only phrase scorer
(score_phrase.py) before any audio existed.
| Stage | Verification |
|---|---|
| BN folding + residual extraction | max err 7.6e-4 vs PyTorch |
| INT8 simulation vs float | 99.3% frame agreement |
| C engine vs INT8 simulation | bit-exact (0.0) |
| C mel frontend vs training frontend | ≤1e-3 log-mel |
| ESP32-S3 step time | 14.2 ms / 40 ms frame (1 core @ 240 MHz) |
engine_c/ is ~300 lines of dependency-free C with one INT8 ring buffer
per layer, so each frame costs only its own ~350k MACs — no window
recomputation. On ESP32 the binding constraint was memory latency, not
math: moving weights out of memory-mapped flash (28.6 → 15.9 ms) and
filling internal SRAM before PSRAM (→ 14.2 ms) mattered far more than
loop optimization.
phoneme_tcn_student.pt deployable model (float, PyTorch state dict)
phoneme_tcn_teacher.pt distillation teacher
model_int8.h INT8 weights + layer table (C header, 340 KB)
act_scales.json activation scales from PTQ calibration
frontend_data.h mel filterbank + Hann window (exact training values)
engine_c/ streaming INT8 engine + decoder + mel frontend (C)
phoneme_engine/ PyTorch model, decoder, enrollment, quantization, scorer
examples/ a universal wake word config (~50 bytes of JSON)
Spot a typed phrase (Python):
import torch
from phoneme_engine.model import PhonemeTCN
from phoneme_engine.features import LogMel
from phoneme_engine.decoder import KeywordSpotter
model = PhonemeTCN().eval()
model.load_state_dict(torch.load("phoneme_tcn_student.pt",
weights_only=True)["model"])
frontend = LogMel().eval()
spotter = KeywordSpotter("hey orbit") # G2P -> HH EY AO R B AH T
with torch.no_grad():
logp = torch.log_softmax(model(frontend(wav)).float(), 2)[0].numpy()
for frame, score in spotter.run(logp):
print(f"detected at {frame * 0.02:.2f}s (score {score:.2f})")
Score a candidate phrase before committing to it:
python score_phrase.py "vucano"
# low reliability: /v/ is acoustically weak and often lost on small mics
A wake word config, in full:
{"phrase": "neptuno",
"spotters": [{"phones": ["N","EH","P","T","UW","N","OW"],
"source": "dictionary", "threshold": -2.5}]}
On a microcontroller: compile engine_c/ with model_int8.h and
frontend_data.h, feed it 40-band mel frames, and pass the logits to
pww_spotter_step(). Wake words are uint8_t arrays of phone ids plus a
float threshold — swap them at runtime.
Apache-2.0. Trained on LibriSpeech (CC BY 4.0) and Common Voice 17 (CC0).
@misc{wakewordswithouttraining2026,
title = {Wake Words Without Training: Open-Vocabulary Wake Word
Creation from Text and a Few Examples},
author = {IOTEverythin},
year = {2026},
url = {https://huggingface.co/IOTEverythin/phoneme-wake-word}
}