Downloads · 30 days
0
MEscriva/gilbert-pyannote-diarization
gilbert-pyannote-diarization is a audio classification model from MEscriva. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for pyannote. The card lists the license as mit.
Model Name: Gilbert Speaker Diarization (v1.0) Model Type: Speaker Diarization Pipeline Base Framework: pyannote.audio 3.x License: MIT Repository: MEscriva/gilbert-pyannote-diarization
Downloads · 30 days
0
Access
Public
Updated Nov 19, 2025
Repo size
—
Likes
1
Public
Click a slice to open those files.
.py32 KB · 75%
From the Hugging Face model README
Model Name: Gilbert Speaker Diarization (v1.0)
Model Type: Speaker Diarization Pipeline
Base Framework: pyannote.audio 3.x
License: MIT
Repository: MEscriva/gilbert-pyannote-diarization
This model provides a speaker diarization pipeline optimized for meeting analysis, built upon the pyannote.audio framework. The implementation includes enhanced post-processing capabilities, overlap detection, and advanced statistical analysis specifically tailored for meeting transcription scenarios. The model is designed to identify and segment speakers in audio recordings with high temporal precision.
The model leverages pre-trained pyannote.audio pipelines, specifically:
pyannote/speaker-diarization-3.1 (default)pyannote/speaker-diarization-community-1, pyannote/speaker-diarization-precision-2The model performance is evaluated using standard diarization metrics:
Based on pyannote.audio benchmarks and internal testing:
| Metric | Performance |
|---|---|
| DER (optimal settings) | < 10% on clean meeting audio |
| Temporal Precision | ± 0.1 seconds |
| Speaker Detection | 95%+ accuracy (known speaker count) |
Note: Performance varies significantly based on audio quality, number of speakers, and overlap frequency.
pip install pyannote.audio pyannote.core torch librosa soundfile
from diarization_pyannote_gilbert import run_gilbert_diarization
results = run_gilbert_diarization(
audio_path="meeting.wav",
model_name="pyannote/speaker-diarization-3.1"
)
# Access results
segments = results["segments"] # Post-processed segments
segments_raw = results["segments_raw"] # Raw pyannote output
overlaps = results["overlaps"] # Detected overlaps
stats = results["stats"] # Per-speaker statistics
# Standard usage (optimal accuracy)
python diarization_pyannote_gilbert.py audio.wav
# With post-processing (improved readability, potential accuracy trade-off)
python diarization_pyannote_gilbert.py audio.wav \
--min-segment 0.5 \
--merge-gaps 0.3
# With known speaker count (improves accuracy)
python diarization_pyannote_gilbert.py audio.wav \
--num_speakers 4
| Parameter | Type | Default | Description |
|---|---|---|---|
model_name | str | pyannote/speaker-diarization-3.1 | Base pyannote model |
num_speakers | int | None | Exact number of speakers (if known) |
min_speakers | int | None | Minimum number of speakers |
max_speakers | int | None | Maximum number of speakers |
min_segment | float | 0.0 | Minimum segment duration (s). 0 = disabled |
merge_gaps | float | 0.0 | Gap threshold for merging (s). 0 = disabled |
use_exclusive | bool | False | Use exclusive speaker diarization |
SPEAKER <file> 1 <start> <duration> <NA> <NA> <speaker_id> <NA> <NA>
[
{
"speaker": "SPEAKER_00",
"start": 0.0,
"end": 3.25
},
...
]
{
"version": "Gilbert-v1.0",
"model": "pyannote/speaker-diarization-3.1",
"num_speakers": 4,
"duration": 3600.0,
"num_segments": 150,
"num_overlaps": 12,
"speaker_stats": {
"SPEAKER_00": {
"total_duration": 900.0,
"num_segments": 45,
"avg_segment_duration": 20.0,
"overlap_duration": 45.2
},
...
}
}
This model is built upon pre-trained pyannote.audio models. The base models were trained on:
Note: This implementation does not include model training; it utilizes pre-trained weights from pyannote.audio.
Evaluation on internal meeting dataset (Gilbert v1 benchmark):
| Dataset | DER (%) | JER (%) | Speakers | Duration (min) |
|---|---|---|---|---|
| Meetings (clean) | 8.5 | 12.3 | 2-4 | 5-60 |
| Meetings (noisy) | 15.2 | 18.7 | 2-4 | 5-60 |
Results may vary based on specific audio characteristics.
If you use this model in your research, please cite:
@software{gilbert_diarization_2024,
title={Gilbert Speaker Diarization Model},
author={MEscriva},
year={2024},
url={https://huggingface.co/MEscriva/gilbert-pyannote-diarization},
version={1.0}
}
This model is released under the MIT License. See LICENSE file for details.
For questions, issues, or contributions, please refer to the repository:
https://huggingface.co/MEscriva/gilbert-pyannote-diarization