Downloads · 30 days
0
sl-data-systems/diarizen-v2-coreml
diarizen-v2-coreml is a machine learning model from sl-data-systems. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for coreml. The card lists the license as cc-by-nc-4.0.
CoreML conversions of the models behind DiariZen's wavlm-large-s80-md-v2 speaker-diarization pipeline, packaged for Apple silicon. Used by the transcriber desktop app through FluidAudio's offline diarization pipeline.…
Downloads · 30 days
0
Access
Public
Updated Sep 22, 2026
Repo size
176 MB
Likes
0
Public
Click a slice to open those files.
.bin176 MB · 99%
From the Hugging Face model README
CoreML conversions of the models behind DiariZen's
wavlm-large-s80-md-v2 speaker-diarization pipeline, packaged for Apple silicon. Used by the
transcriber desktop app through FluidAudio's
offline diarization pipeline. Every conversion is an exact algebraic rewrite of the source
checkpoint (reshapes, framing, normalization identities, activation identities); nothing was
retrained, fine-tuned or further pruned.
| file | source checkpoint | interface | placement |
|---|---|---|---|
DiariZenSegmentation.mlmodelc | BUT-FIT/diarizen-wavlm-large-s80-md-v2 (pruned WavLM-Large + Conformer, 4 local speakers, 16-class powerset) | waveform (1, 256000) float32, 16 kHz mono → logprobs (1, 799, 16) | Neural Engine (1253 of 1255 ops) |
DiariZenFBank.mlmodelc | Kaldi fbank + per-window mean normalization as used by pyannote's WeSpeaker wrapper | audio (B, 1, 256000), B in 1…32 → fbank_features (B, 1, 80, 1598) | CPU |
DiariZenEmbedding.mlmodelc | pyannote/wespeaker-voxceleb-resnet34-LM | fbank_features (1, 1, 80, 1598) + weights (4, 799) mask lanes → embedding (4, 256) | Neural Engine |
PldaRho.mlmodelc, plda-parameters.json, xvector-transform.json | DiariZen's PLDA (plda.npz, xvec_transform.npz), byte-identical to pyannote community-1's | 256 → 128 | CPU |
The embedding model runs the ResNet trunk once per 16 s window and pools up to four speaker masks from it, which is what DiariZen's per-slot embedding computes, at a quarter of the calls.
Verified against the PyTorch pipeline on ICSI meetings: segmentation log-probabilities within fp16 rounding (argmax agreement ≥ 99.75 %), embeddings cosine ≥ 0.99997, and identical per-embedding cluster labels through PLDA + AHC + VBx; meeting-level DER equal to the reference.
Requirements: macOS 14 or later, Apple silicon. First load compiles the segmentation model for the Neural Engine (10–25 s, cached afterwards).
pyannote/wespeaker-voxceleb-resnet34-LM
(CC BY 4.0), itself a port of WeSpeaker's
VoxCeleb ResNet34-LM.Please cite DiariZen if you use these models:
Jiangyu Han et al., "Leveraging Self-Supervised Learning for Speaker Diarization" and "Fine-tuning Self-Supervised Models for Speaker Diarization with Structured Pruning" (BUT-FIT, 2025).