Downloads · 30 days
29
33% of all-time downloads
aufklarer/Ultra-Sortformer-Diarization-CoreML
Ultra-Sortformer-Diarization-CoreML is a machine learning model from aufklarer. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for coreml. The card lists the license as apache-2.0.
CoreML conversion of devsy0117/ultradiarstreamingsortformer8spkv1, the Ultra-Sortformer fine-tune that widens NVIDIA's streaming Sortformer speaker head from four to eight speakers with SVD-initialized rows and split-…
Downloads · 30 days
29
33% of all-time downloads
All-time downloads
88
Public
Repo size
236 MB
Likes
1
Public
Click a slice to open those files.
.bin470 MB · 100%
From the Hugging Face model README
CoreML conversion of
devsy0117/ultra_diar_streaming_sortformer_8spk_v1,
the Ultra-Sortformer
fine-tune that widens NVIDIA's streaming Sortformer speaker head from four to
eight speakers with SVD-initialized rows and split-learning-rate training.
Same conversion pipeline, streaming state protocol, and Arrival-Order Speaker
Cache host algorithm as
aufklarer/Sortformer-Diarization-CoreML
(the 4-speaker base). Only the head width and the emitted probability tensor
change: speaker_preds [1, 242, 8] (188 speaker-cache + 40 FIFO + 14 chunk
frames), with the base model's cache geometry preserved.
| File | Purpose |
|---|---|
Sortformer_streaming.mlmodelc | Compiled streaming variant — 480 ms of new audio per call, cache state crosses the interface as tensors |
Sortformer_streaming.mlpackage | Source package for recompilation |
config_streaming.json | Loader configuration (num_speakers: 8, chunk/cache dims) |
The exported step was driven by NeMo's own streaming_feat_loader and
streaming_update_async and matched the checkpoint's native
forward_streaming loop (streaming-parity pytest in the conversion pipeline,
passing).
nvidia/diar_streaming_sortformer_4spk-v2.1
(NVIDIA Open Model License).devsy0117/ultra_diar_streaming_sortformer_8spk_v1
(Ultra-Sortformer,
Apache-2.0), trained on synthetic multi-speaker sessions built from the
AI Hub multi-speaker speech synthesis corpus (Korean).Note: the upstream project publishes real-corpus rankings only for the 4-speaker base model. Evaluate this 8-speaker variant on your own data before preferring it — the wider head targets crowded scenes, and behavior on 2–4 speaker audio should be regression-checked per application.