Downloads · 30 days
19
26% of all-time downloads
hgsilver911/speaker-diarization-coreml
speaker-diarization-coreml is a voice activity detection model from hgsilver911. Use it for the voice activity detection task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
[](https://discord.gg/WNsvaCtmDe) [](https://github.com/FluidInference/FluidAudio)
Downloads · 30 days
19
26% of all-time downloads
All-time downloads
72
Public
Repo size
254 MB
Likes
0
Public
Click a slice to open those files.
.bin123 MB · 95%
From the Hugging Face model README
Speaker diarization based on pyannote models optimized for Apple Neural Engine.
Models are trained on acoustic signatures so it supports any lanugage.
See the SDK for more details https://github.com/FluidInference/FluidAudio
Please note that the SDK itself is Apache 2.0, but the parent model from Pyannote is cc-by-4.0
See the origianl model for detailed DER benchmark, for the purpose of our conversion, we tried to match the original model as much as possible:
The models on CoreML exhibit a ~10x Speedup on CPU and ~20x speed up on GPU.

Due to different precisions, there are minor differences in the values generated but the differences are mostly negilible, though it does account for some errors that needs to be adjusted during clustering:

We see this when running the end to end pipeline with the Pytorch model versus the Core ML model (patched the Pyannote pipeline to run the Core ML model instead). The DER and JER is ~1% compared to the Pytorch model as we're dropping the precision to fp32

@inproceedings{Plaquet23,
author={Alexis Plaquet and Hervé Bredin},
title={{Powerset multi-class cross entropy loss for neural speaker diarization}},
year=2023,
booktitle={Proc. INTERSPEECH 2023},
}
@inproceedings{Wang2023,
title={Wespeaker: A research and production oriented speaker embedding learning toolkit},
author={Wang, Hongji and Liang, Chengdong and Wang, Shuai and Chen, Zhengyang and Zhang, Binbin and Xiang, Xu and Deng, Yanlei and Qian, Yanmin},
booktitle={ICASSP 2023, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
pages={1--5},
year={2023},
organization={IEEE}
}
@article{Landini2022,
author={Landini, Federico and Profant, J{\'a}n and Diez, Mireia and Burget, Luk{\'a}{\v{s}}},
title={{Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks}},
year={2022},
journal={Computer Speech \& Language},
}