Downloads · 30 days
1K
100% of all-time downloads
mlx-community/Nemotron-3-Diarization
Nemotron-3-Diarization is a voice activity detection model from mlx-community. Use it for the voice activity detection task on the model card, and read the license before you ship it in a product. It is set up for mlx-audio. The card lists the license as openmdw-1.1.
Converted from nvidia/Nemotron-3-Diarization for use with mlx-audio.
Downloads · 30 days
1K
100% of all-time downloads
All-time downloads
1K
Public
Parameters
99.3M
199 MB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors199 MB · 100%
How the weights are stored.
BF1699.2M · 100%
From the Hugging Face model README
Converted from nvidia/Nemotron-3-Diarization for use with mlx-audio.
It predicts up to eight speakers at 10 ms resolution from 16 kHz mono audio.
from mlx_audio.vad import load
model = load("mlx-community/Nemotron-3-Diarization", strict=True)
result = model.generate("meeting.wav")
print(result.text)
# Incremental results retain speaker identities through the AOSC and FIFO.
for result in model.generate_stream("meeting.wav"):
for segment in result.segments:
print(segment.start, segment.end, segment.speaker)
For live PCM input, call model.feed(chunk, state) with a state from
model.init_streaming_state() and 16 kHz mono chunks. Flush the final partial
chunk and lookahead with model.feed([], state, final=True). Timestamps are
absolute within the recording. Labels are generic arrival-order speaker IDs;
the model does not identify people. Overlapping speakers may be active together.
Consult the upstream model card for training data, evaluation and license information.