Downloads · 30 days
0
0% of all-time downloads
audarai/Audar-Diarization-V1
Audar-Diarization-V1 is a audio classification model from audarai. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for nemo. The card lists the license as other.
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
28
Public
Parameters
118M
471 MB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors471 MB · 100%
From the Hugging Face model README
From Arabic to the world.
<p><a href="#-what-it-is"><b>🧭 Overview</b></a> · <a href="#-benchmarks"><b>📊 Benchmarks</b></a> · <a href="#-quickstart"><b>⚡ Quickstart</b></a> · <a href="#-real-time-streaming"><b>🎙️ Streaming</b></a> · <a href="#-files"><b>📦 Files</b></a> · <a href="https://github.com/AudarAI/Audar-diarization-V1"><b>🐙 GitHub</b></a> · <a href="https://www.audarai.com"><b>☁️ Audar API</b></a> · <a href="https://www.audarai.com/license/audarai-community-license-v1.0/"><b>📜 License</b></a></p> </div>Audar-Diarization-V1 answers "who spoke when" — in real time, for up to 8 speakers, across hour-long multi-speaker audio. It is the speaker-attribution engine of the Audar realtime stack: paired with Audar-ASR-V1 it turns a verbatim transcript into a speaker-labeled one — the difference between an undifferentiated wall of text and a minutes-ready board record.
It is built on NVIDIA's Streaming Sortformer v2.1 and advanced in-house through Audar's diarization program:
The result streams on a single GPU with 1.04 s algorithmic latency and a 0.003 real-time factor (1 s of audio processed in ~3 ms), while posting the lowest DER of any evaluated system on all eight benchmark corpora.
Evaluated with dscore at a 0.25 s collar, ignoring overlap (DIHARD protocol) on the official
dev/eval splits of 8 corpora spanning meetings, dinner parties, broadcast, and in-the-wild audio.
Audar-Diarization-V1 posts the lowest DER on every corpus and a macro DER of 22.03 % — beating
pyannote 3.1 by 7.63 pp and stock Sortformer v2.1 by 12.16 pp.
| System | AMI | AliMeeting | DiPCo | ICSI | MSDWild-few | MSDWild-many | VoxConverse | CHiME-6 | Macro |
|---|---|---|---|---|---|---|---|---|---|
| Audar-Diarization-V1 | 15.24 | 18.70 | 23.77 | 14.46 | 21.09 | 29.41 | 8.55 | 45.00 | 22.03 |
| pyannote 3.1 | 28.60 | 27.38 | 30.72 | 22.48 | 27.12 | 34.83 | 12.92 | 53.19 | 29.66 |
| Sortformer v2.1 | 24.84 | 25.94 | 33.80 | 23.22 | 36.92 | 50.77 | 17.06 | 60.97 | 34.19 |
| System | Miss | False alarm | Confusion | DER |
|---|---|---|---|---|
| Audar-Diarization-V1 | 10.33 | 6.90 | 4.80 | 22.03 |
| pyannote 3.1 | 8.72 | 3.38 | 17.56 | 29.66 |
| Sortformer v2.1 | 11.91 | 5.04 | 17.24 | 34.19 |
The advantage is confusion: 4.80 % vs 17.2–17.6 % — a 3.6× reduction, from the AOSC's stable identity tracking. The slightly higher false-alarm rate reflects a deliberately assertive streaming VAD (a missed utterance costs more than a brief false activation in live transcription) and is tunable via the onset threshold.
| Audar-Diarization-V1 | Sortformer v2.1 | pyannote 3.1 |
|---|---|---|
| 10.29 | 12.22 | 18.51 |
Identity also holds on the longest sessions in the benchmark — e.g. a 74-minute, 5-speaker ICSI meeting at 22.2 % DER with ~2 % confusion.
Ships as a single fp32 safetensors bundle — model.safetensors + config.yaml + load_diarizer.py.
The loader instantiates the NeMo Sortformer model and loads the weights directly (no .nemo tar):
# needs: nemo_toolkit[asr]>=2.6, safetensors, omegaconf
from huggingface_hub import snapshot_download
import sys; sys.path.insert(0, snapshot_download("audarai/Audar-Diarization-V1"))
from load_diarizer import load_diarizer
model = load_diarizer() # fp32, CUDA (device="cpu" also works)
segs = model.diarize(audio=["meeting.wav"], batch_size=1)
# → RTTM-style [(start_s, end_s, speaker_slot), ...] per file
Runnable inference examples (offline diarization + RTTM export), the full 8-corpus benchmark, and reproduction steps are open at github.com/AudarAI/Audar-diarization-V1.
The same checkpoint runs true streaming via forward_streaming_step with persistent spkcache / FIFO
state: ~1 s chunks, 80 ms prediction frames, up to 8 concurrent speakers, and session-stable slot
labels that never rewrite once committed. Algorithmic latency is 1.04 s at a 0.003 real-time
factor on a single GPU.
Speaker-attributed transcription. The Audar serving gateway runs diarization in parallel with Audar-ASR-V1 and assigns each transcribed word to the speaker dominant during its time span — so combined latency is the max of the two streams, not the sum. One deployment exposes ASR-only, diarization-only, and ASR+diarization endpoints over HTTP and an OpenAI-Realtime-compatible WebSocket. For a managed, production-hosted endpoint, see the Audar API.
| File | What it is |
|---|---|
model.safetensors | fp32 weights — bit-exact, lossless, safetensors (safe, zero-copy mmap, no pickle) |
config.yaml | model config (the .nemo's model_config.yaml) |
load_diarizer.py | self-contained loader (instantiates the NeMo model + loads the weights) |
Lossless fp32, and faster to load. The model.safetensors carries the full-precision weights
bit-for-bit — verified by a round-trip check (990/990 tensors identical) and by downcasting to the prior
fp16 release with zero mismatches across all 971 float tensors, so it reproduces the exact model. The
weight-load step is ~28× faster than the legacy .nemo (≈12 ms mmap vs ≈344 ms untar + unpickle),
and safetensors is the safe, community-standard format (no arbitrary-code pickle path).
load_diarizer.py auto-detects weight dtype, so you can quantize model.safetensors to fp16 yourself and
it will load unchanged. When fp16 is detected the loader keeps the preprocessor (STFT/mel) in fp32 and runs
the streaming path under torch.set_default_dtype(torch.float16) (NeMo's streaming state is otherwise
created dtype-less). ONNX export is not supported out-of-the-box (NeMo 2.6.2's Sortformer export needs
streaming-state wiring).
Intended use. Speaker-attributed meeting/broadcast/call-center transcription, board and panel recordings, and any real-time or offline "who spoke when" task — cloud, on-prem, or edge.
Limitations.
Released under the AudarAI Community License v1.0 — research and limited commercial use for qualifying Community Entities; enterprise, large-scale, or model-as-a-service use requires an AudarAI Enterprise License. See audarai.com/license/audarai-community-license-v1.0, or contact [email protected] for enterprise licensing.
@techreport{audar-diarization-v1-2026,
title = {Audar-Diarization-V1: Real-Time Streaming Speaker Diarization for Long-Form, Multi-Speaker Audio},
author = {Audar AI Team},
institution = {AudarAI},
year = {2026},
url = {https://huggingface.co/audarai/Audar-Diarization-V1}
}
AudarAI starts with Arabic — and expands to the world.
</div>We are building advanced multilingual audio intelligence that helps individuals, enterprises, and governments communicate across languages, cultures, and borders. By combining Arabic-first speech technology with global multilingual AI, AudarAI transforms voice into understanding, interaction, and connection.
Our work spans speech recognition, speech understanding, speaker diarization, voice-enabled digital assistants, human-computer interaction, and intelligent audio systems designed for real-world impact. From empowering people to access technology in their native language to helping organizations communicate globally, AudarAI is shaping a future where every voice can be heard, understood, and connected.
Arabic-first. Multilingual by design. Human-centered at heart.
<div align="center">🌐 www.audarai.com · 🤗 Hugging Face · GitHub · [email protected]
© 2026 AUDARAI PTE. LTD. · Licensed under the AudarAI Community License v1.0
</div>