Downloads · 30 days
0
AswanthCManoj/moonshine-eou-detector
moonshine-eou-detector is a machine learning model from AswanthCManoj. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
End-of-utterance (EOU) confirmation gate for streaming voice pipelines — runs alongside Silero VAD to decide when a caller's turn is truly finished.
Downloads · 30 days
0
Access
Public
Updated Sep 6, 2026
Repo size
36.9 MB
Likes
0
Public
Click a slice to open those files.
.data31.4 MB · 85%
From the Hugging Face model README
End-of-utterance (EOU) confirmation gate for streaming voice pipelines — runs alongside Silero VAD to decide when a caller's turn is truly finished.
All-ONNX inference, no PyTorch at runtime.
| File | Size | Description |
|---|---|---|
moonshine_enc_fp32.onnx | 0.9 MB | Moonshine-tiny encoder graph (7.7M params, 288D hidden) |
moonshine_enc_fp32.onnx.data | 30 MB | Encoder weights (external data) |
eou_head.onnx | 2.2 MB | Trained binary classification head |
silero_vad.onnx | 2.2 MB | Silero VAD v5 (bundled for convenience) |
import numpy as np
import onnxruntime as ort
# Load sessions
enc = ort.InferenceSession("moonshine_enc_fp32.onnx")
head = ort.InferenceSession("eou_head.onnx")
# Predict P(turn complete) from 16 kHz float32 audio
wav = np.zeros(16000, dtype=np.float32).reshape(1, -1) # 1 second
features = enc.run(None, {"input_values": wav})[0]
frame_len = np.array([features.shape[1]], dtype=np.int64)
logit = head.run(None, {"encoder_features": features, "frame_lengths": frame_len})[0]
prob = float(1.0 / (1.0 + np.exp(-logit[0])))
The model is consulted at up to 4 checkpoints during caller silence, each with a relaxing gate:
| Silence | Threshold | Rationale |
|---|---|---|
| 96 ms | 0.85 | Very confident — fast commit |
| 192 ms | 0.65 | Moderate confidence |
| 320 ms | model default (0.50) | Calibrated threshold |
| 480 ms | 0.40 | Lenient — long pause likely means done |
| 576 ms | (timeout) | Hard backstop, commit regardless |
The bundled silero_vad.onnx requires 64 samples of context prepended to each 512-sample window (576 total input). Without this, the model produces near-zero probabilities and speech detection silently fails.
The official Silero OnnxWrapper and PyTorch model handle this internally. Custom ONNX code must replicate it:
_CONTEXT_SIZE = 64
context = np.zeros(64, dtype=np.float32)
def process_chunk(chunk_512):
x = np.concatenate([context, chunk_512.astype(np.float32)])
context[:] = x[-_CONTEXT_SIZE:]
out, state = session.run(None, {
"input": x.reshape(1, -1),
"state": state,
"sr": np.array(16000, dtype=np.int64),
})
return float(out[0, 0])
(1, seq) float32(2, 1, 128) (~1 KB) + 64-sample context buffer