Downloads · 30 days
0
Transcripta/sherpa-onnx-speaker-diarization-wasm
sherpa-onnx-speaker-diarization-wasm is a machine learning model from Transcripta. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Prebuilt WebAssembly assets for fully on-device speaker diarization ("who spoke when"), redistributed for the TranscriptAI / FileWhirl on-device audio features so they can be fetched with CORS at runtime and cached in…
Downloads · 30 days
0
Access
Public
Updated Sep 6, 2026
Repo size
58 MB
Likes
0
Public
Click a slice to open those files.
.data45.6 MB · 79%
From the Hugging Face model README
Prebuilt WebAssembly assets for fully on-device speaker diarization ("who spoke when"), redistributed for the TranscriptAI / FileWhirl on-device audio features so they can be fetched with CORS at runtime and cached in the browser. No audio leaves the device.
These are the unmodified prebuilt artifacts from sherpa-onnx v1.13.7 (k2-fsa),
speaker-diarization WASM target. The pipeline is pyannote-segmentation-3.0 (ONNX) +
a speaker-embedding model + agglomerative clustering, all inside the wasm module. The
models are baked into the .data preload.
sherpa-onnx-wasm-main-speaker-diarization.js / .wasm / .data — the Emscripten buildsherpa-onnx-speaker-diarization.js — the JS API wrapper (createOfflineSpeakerDiarization)Load the three scripts (define Module = {} with onRuntimeInitialized first), then:
const sd = createOfflineSpeakerDiarization(Module); const segments = sd.process(float32Mono16k);
Each segment is { start, end, speaker } (seconds, integer cluster id).
Redistributed under the above terms with attribution.