Downloads · 30 days
0
AswanthCManoj/moonshine-eou-head
moonshine-eou-head is a audio classification model from AswanthCManoj. Use it for the audio classification task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A lightweight end-of-utterance (EOU) confirmation head trained on top of Moonshine-tiny encoder features. Designed to run alongside Silero VAD as a secondary contextual confirmer — when VAD detects 50-80ms of silence,…
Downloads · 30 days
0
Access
Public
Updated Sep 6, 2026
Repo size
39.2 MB
Likes
0
Public
Click a slice to open those files.
.data31.4 MB · 80%
From the Hugging Face model README
A lightweight end-of-utterance (EOU) confirmation head trained on top of Moonshine-tiny encoder features. Designed to run alongside Silero VAD as a secondary contextual confirmer — when VAD detects 50-80ms of silence, this model confirms whether the speaker is actually done talking or just pausing.
Linear(288 → 128) + GELUBiGRU(128, 2 layers, bidirectional)Attention pooling over GRU outputsLayerNorm(256) → Linear(256 → 128) + GELU → Linear(128 → 1)Trained on the full SmartTurn v3.2 dataset (271K clips, 50/50 balanced):
| Metric | Value |
|---|---|
| AUC | 0.9665 |
| Accuracy | 0.908 |
| F1 | 0.910 |
Full pipeline is 100% ONNX Runtime — no PyTorch needed at inference.
| Component | 3s audio |
|---|---|
| Silero VAD (per 32ms chunk) | ~1ms |
| Moonshine encoder (ONNX fp32) | ~25ms |
| EOU head (ONNX) | ~1ms |
| Total EOU decision | ~26ms |
Clone and run — no pip install needed, just numpy and onnxruntime:
from eou_gate import create_engine, Event
# Once per process — loads all 3 ONNX models
engine = create_engine(model_dir="models")
# Per call / user — lightweight state only
stream = engine.create_stream(min_silence_ms=60)
# Feed 16kHz float32 audio chunks (any size)
for chunk in audio_chunks:
event = stream.process_chunk(chunk)
if event == Event.TURN_END:
print("User finished speaking")
elif event == Event.PAUSE:
print("User is pausing (not done yet)")
# Or async
event = await stream.aprocess_chunk(chunk)
# Direct EOU probability (no VAD, call on demand)
prob = engine.eou_infer(audio_3s) # -> float in [0, 1]
The engine is a shared singleton; each stream holds only its own VAD state and ring buffer. Thread pool handles concurrent inference:
import asyncio
engine = create_engine(model_dir="models")
async def handle_call(user_audio_stream):
stream = engine.create_stream(min_silence_ms=60)
async for chunk in user_audio_stream:
event = await stream.aprocess_chunk(chunk)
if event == Event.TURN_END:
# respond to user
stream.reset()
# 8 concurrent calls sharing one engine
await asyncio.gather(*[handle_call(s) for s in streams])
import onnxruntime as ort
import numpy as np
enc = ort.InferenceSession("models/moonshine_enc_fp32.onnx")
head = ort.InferenceSession("models/eou_head.onnx")
# Raw 16kHz float32 audio → EOU probability
audio = load_audio("clip.wav") # (samples,) float32
features = enc.run(None, {"input_values": audio.reshape(1, -1)})[0]
logit = head.run(None, {
"encoder_features": features,
"frame_lengths": np.array([features.shape[1]], dtype=np.int64),
})[0]
prob = 1.0 / (1.0 + np.exp(-logit[0]))
# prob > 0.5 → turn is done
models/
├── silero_vad.onnx # Silero VAD v5 (2.2 MB)
├── moonshine_enc_fp32.onnx # Moonshine-tiny encoder ONNX (0.9 MB + 30 MB data)
├── moonshine_enc_fp32.onnx.data # Encoder weights (external data)
└── eou_head.onnx # Trained EOU head ONNX (2.2 MB)
eou_gate/ # Drop-in Python package
├── __init__.py # create_engine() factory
├── engine.py # EOUEngine — loads & runs all models
├── stream.py # StreamState — per-user VAD + EOU state
└── audio.py # RingBuffer — circular audio buffer
train.py # Full training script (Colab/RunPod ready)
eou_head_best.pt # PyTorch checkpoint (for fine-tuning)
min_silence_ms (e.g. 60ms) after speech:TURN_END; if not → PAUSE (keep listening)This lets min_silence_ms drop from ~256ms to ~60-80ms — making voice agents feel much more responsive without cutting people off mid-sentence.
Only two runtime dependencies — no PyTorch or transformers needed:
numpy
onnxruntime
This model is a confirmation gate, not a standalone VAD. It is designed to:
min_silence_ms to drop from ~256ms to ~60-80ms for faster turn-taking in voice agentsMIT