Downloads · 30 days
166
11% of all-time downloads
TheStageAI/silero-vad
silero-vad is a voice activity detection model from TheStageAI. Use it for the voice activity detection task on the model card, and read the license before you ship it in a product. It is set up for coreml. The card lists the license as mit.
Downloads · 30 days
166
11% of all-time downloads
All-time downloads
1.5K
Public
Repo size
11.8 MB
Likes
1
Public
Click a slice to open those files.
.zip2 MB · 52%
From the Hugging Face model README

Original model: snakers4/silero-vad (Silero Team)
On-device Silero VAD packaged by TheStage AI for Apple Silicon. Stateful per-chunk speech probability — used as the mic gate in voice agents and as Whisper's optional internal VAD pre-pass.
| Runtime | TheStage Apple SDK (SileroVAD / start_model) |
| Chunk | 512 samples (32 ms @ 16 kHz) |
| HF engines | TheStageAI/silero-vad |
This repo hosts CoreML engine bundles for Silero VAD. The model keeps an LSTM hidden state across calls — reset between independent utterances.
| Property | Value |
|---|---|
| Hardware | Apple Silicon Mac, or physical iPhone / iPad |
| macOS | 15.0+ |
| iOS | 18.0+ |
| Xcode | 16.0+ |
| Swift | 6.0+ |
| Flutter (optional) | 3.24+ |
Simulator is not supported — run on real Apple Silicon hardware.
Create a token at app.thestage.ai and pass it to the SDK:
try await TheStageAI.shared.initialize(apiToken: "th_…")
await TheStageFlutterSDK.initialize(api_token: 'th_…');
Token is checked online in initialize (once per app process when reachable). Inference runs fully on-device. Offline initialize fails — reconnect and call initialize again.
Docs: TheStage Apple SDK · VAD
In Xcode: File → Add Package Dependencies…, paste https://github.com/TheStageAI/AppleSDK.git, and add the TheStageSDK product. Or in Package.swift:
.package(
url: "https://github.com/TheStageAI/AppleSDK.git",
exact: Version(1, 1, 0)
)
# pubspec.yaml
dependencies:
thestage_apple_sdk:
git:
url: https://github.com/TheStageAI/AppleSDK.git
path: plugin/thestage_apple_sdk
ref: 1.1.0
import TheStageSDK
try await TheStageAI.shared.initialize(apiToken: "th_…")
try await TheStageAI.shared.start_model(
model_name: "vad",
engines_path: "TheStageAI/silero-vad"
)
// Exactly 512 samples @ 16 kHz mono per call
let result = try TheStageAI.shared.infer(
model_name: "vad",
input_json: ["audio": audio_chunk]
)
let probability = result[0]["probability"] as! Double
if probability > 0.5 {
print("Speech detected")
}
Pass "reset_state": true between independent utterances.
await TheStageFlutterSDK.start_model(
model_name: 'vad',
engines_path: 'TheStageAI/silero-vad',
);
final result = await TheStageFlutterSDK.infer(
model_name: 'vad',
input_json: {'audio': audio_chunk}, // Float32List, length 512
);
final probability = result[0]['probability'] as double;
| Sample rate | 16 kHz mono Float |
| Chunk size | exactly 512 samples (smaller chunks are zero-padded; larger rejected) |
| Output | probability in [0.0, 1.0] |
| State | Stateful LSTM — reset between utterances |
This work builds on Silero VAD by Silero Team: snakers4/silero-vad.
This Apple Silicon package is produced by TheStage AI; it is not an official Silero release.