Downloads · 30 days
11
19% of all-time downloads
AXERA-TECH/DiariZen
DiariZen is a voice activity detection model from AXERA-TECH. Use it for the voice activity detection task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as cc-by-nc-4.0.
CPU+NPU hybrid speaker diarization segmentation for AX650 NPU.
Downloads · 30 days
11
19% of all-time downloads
All-time downloads
58
Public
Repo size
276 MB
Likes
1
Public
Click a slice to open those files.
.onnx275 MB · 99%
From the Hugging Face model README
CPU+NPU hybrid speaker diarization segmentation for AX650 NPU.
Audio (16kHz mono, any length)
→ CPU: resample → 4s sliding window → LayerNorm
→ AX650 NPU: CNN feature extractor (7 conv, U16, 17.7ms)
→ CPU: WavLM Transformer (24L) + Conformer (4L) + Classifier (251ms)
→ Log-probabilities (1, 199, 11) per window
| Stage | Time | Hardware |
|---|---|---|
| CNN | 17.7 ms | AX650 NPU @1GHz |
| Backend | 251 ms | CPU ONNX Runtime |
| Total | 269 ms | 14.9x real-time |
End-to-end cosine: 0.9997 vs full FP32 reference.
models/ cnn_features.axmodel + backend.onnx
python/ Python SDK (diarizen_sdk)
cpp/ C++ SDK (diarizen_segmenter)
model_convert/ ONNX export + Pulsar2 compile config
reports/ SDK and simulation reports
pip install -r python/requirements.txt
python python/diarizen_sdk/example.py audio.wav \
--cnn-model models/cnn_features.axmodel \
--backend-model models/backend.onnx