Downloads · 30 days
31
41% of all-time downloads
soniqo/Sortformer-Diarization-4spk-ONNX
Sortformer-Diarization-4spk-ONNX is a machine learning model from soniqo. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for onnx. The card lists the license as other.
ONNX export of NVIDIA's streaming Sortformer 4-speaker diarization model, for hosts that drive ONNX Runtime. Apple platforms have a CoreML build; this is the same checkpoint for everything else.
Downloads · 30 days
31
41% of all-time downloads
All-time downloads
76
Public
Repo size
953 MB
Likes
0
Public
Click a slice to open those files.
.onnx477 MB · 50%
From the Hugging Face model README
ONNX export of NVIDIA's streaming Sortformer 4-speaker diarization model, for hosts that drive ONNX Runtime. Apple platforms have a CoreML build; this is the same checkpoint for everything else.
Exported from nvidia/diar_streaming_sortformer_4spk-v2.1 with no retraining
and no change to weights, architecture or configuration. The export composes
the model's pre-encoder and head into one graph per streaming step and traces
it at the default variant's static shapes.
in: chunk[1, 3048, 128] mel features, 128 bins
chunk_lengths[1]
spkcache[1, 188, 512] speaker cache
spkcache_lengths[1]
fifo[1, 40, 512]
fifo_lengths[1]
out: spkcache_fifo_chunk_preds[1, 609, 4] per-frame activity, 4 speakers
chunk_pre_encode_embs[1, 381, 512]
chunk_pre_encode_lengths[1]
One call consumes ~30 seconds of new audio and needs ~3.2 seconds of audio after the chunk it reports on.
The graph takes spkcache and fifo as inputs and returns embeddings; it does
not update them. Deciding what enters the cache, what waits in the FIFO and
what is evicted is the caller's job, and that bookkeeping — the Arrival-Order
Speaker Cache — is what makes a speaker index mean the same person for a whole
recording. A host that skips it gets per-call speaker numbering and none of the
model's advantage over window-local segmenters.
Eviction is not FIFO. Frames are scored by how confidently one speaker and no
other is active, boosted twice so every speaker keeps a minimum share and none
dominates, then the highest-scoring 188 are kept in chronological order with
unused slots filled by a running mean-silence embedding. NeMo's
SortformerModules.streaming_update_async is the reference.
The cache geometry above belongs to this export. Pairing it with another variant's update period evicts the wrong frames while looking healthy.
Against NeMo's own forward_streaming on 132 seconds of two-speaker audio,
with the speaker cache overflowing:
| mean absolute error | 0.0149 |
| decision agreement | 0.9850 |
| throughput | 42.4x realtime, 5 calls, 615 ms each |
A bare ONNX Runtime session, without the reference model loaded beside it, is 505 MB resident and 1251 MB peak, and one call takes 467 ms to cover 30 s of new audio — 64x realtime at 12 intra-op threads. The 42.4x above comes from the parity harness, which carries PyTorch and NeMo in the same process; use it to compare against the reference, not to size a deployment.
Measured on a 24-core x86 CPU through ONNX Runtime's CPU provider. The same loop driven by PyTorch instead of this graph gives identical numbers, so the export contributes no error of its own — the difference is between NeMo's two internal streaming paths.
It tracks at most four speakers and degrades beyond that.
sortformer-default.onnx and config.json, which carries the cache sizes and
chunk geometry a host needs to size its buffers.