Downloads · 30 days
0
shareed2k/callmind-pyannote-segmentation
callmind-pyannote-segmentation is a voice activity detection model from shareed2k. Use it for the voice activity detection task on the model card, and read the license before you ship it in a product. It is set up for onnx. The card lists the license as mit.
pyannote/segmentation-3.0 re-exported so that it can be loaded by tract, the pure-Rust ONNX runtime. Functionally identical to the upstream export; only the graph's shape handling differs.
Downloads · 30 days
0
Access
Public
Updated Aug 23, 2026
Repo size
6 MB
Likes
0
Public
Click a slice to open those files.
.onnx6 MB · 100%
From the Hugging Face model README
pyannote/segmentation-3.0 re-exported so that it can be loaded by
tract, the pure-Rust ONNX runtime. Functionally
identical to the upstream export; only the graph's shape handling differs.
Used by CallMind to measure how many speakers a recording holds instead of assuming two.
The published ONNX export cannot be loaded by tract. It contains an If node
guarding a branch on the input's shape, plus symbolic dimension arithmetic that
tract declines to prove equal (-24 + T/10 against -25 + (T+9)/10). Both
disappear once the input is fixed to a single 10-second chunk: the branch
condition becomes a constant and ordinary constant folding removes the node.
Verified against the source export:
| check | result |
|---|---|
| output change from folding | exactly 0.0 |
tract against onnxruntime | 6.7e-4 max, 3.0e-4 mean (f32 accumulation) |
| per-frame decision agreement | 589 / 589 frames |
Operators after folding are all standard ai.onnx: InstanceNormalization,
Conv, MaxPool, LeakyRelu, Transpose, LSTM, Reshape, Gemm,
LogSoftmax, Abs.
The int8 variant is deliberately not provided: it folds to a graph containing
DynamicQuantizeLSTM from the com.microsoft domain, which tract does not
implement.
input x | [1, 1, 160000] — one 10-second chunk, 16 kHz mono, float32 |
output y | [1, 589, 7] — per-frame log-probabilities over a speaker powerset |
The seven classes are ∅, {1}, {2}, {3}, {1,2}, {1,3}, {2,3} — up to three
speakers with the two-at-once combinations, which is how overlapping speech is
represented. Frame duration is 10000 / 589 ≈ 16.98 ms.
Longer audio is processed one chunk at a time; the final chunk is zero-padded.
Against labelled recordings — four single-speaker recordings confirmed by their owner and thirty two-party phone calls — taking the median number of distinct speakers seen per chunk:
| statistic | single speaker | two-party |
|---|---|---|
| maximum per chunk | 4/4 | 20/30 |
| median per chunk | 4/4 | 26/30 |
The maximum is the wrong statistic despite looking natural: it is the maximum of a noisy quantity, so it grows with recording length and long calls reliably report one speaker too many.
python3 -m venv .venv && .venv/bin/pip install onnx onnxruntime numpy
.venv/bin/python scripts/export_pyannote_segmentation.py \
--out models/diarization/segmentation.onnx
The script re-downloads the source, applies the transformation and refuses to
write a result whose output drifted, whose If node survived, or which contains
non-standard operator domains.
MIT, inherited unchanged from upstream.
pyannote/segmentation-3.0, Copyright (c) 2022 CNRS.csukuangfj/sherpa-onnx-pyannote-segmentation-3-0.ivrit-ai/pyannote-segmentation-3.0.If you use pyannote in research, cite the upstream work:
@inproceedings{Plaquet23,
author={Alexis Plaquet and Hervé Bredin},
title={{Powerset multi-class cross entropy loss for neural speaker diarization}},
year=2023,
booktitle={Proc. INTERSPEECH 2023},
}
@inproceedings{Bredin23,
author={Hervé Bredin},
title={{pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe}},
year=2023,
booktitle={Proc. INTERSPEECH 2023},
}