Downloads · 30 days
26
15% of all-time downloads
GradientDescent2718/LS-EEND-ONNX
LS-EEND-ONNX is a audio classification model from GradientDescent2718. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for onnx. The card lists the license as mit.
ONNX exports of LS-EEND, a long-form streaming end-to-end neural diarization model with online attractor extraction.
Downloads · 30 days
26
15% of all-time downloads
All-time downloads
174
Public
Repo size
178 MB
Likes
1
Public
Click a slice to open those files.
.onnx178 MB · 100%
From the Hugging Face model README
ONNX exports of LS-EEND, a long-form streaming end-to-end neural diarization model with online attractor extraction.
This repository contains non-quantized ONNX step models for four LS-EEND variants:
AMICALLHOMEDIHARD IIDIHARD IIIThese models are intended for stateful streaming inference. Each package runs one LS-EEND step at a time with explicit recurrent/cache tensors, rather than processing an entire utterance in a single call.
Each variant directory contains:
*.onnx: the ONNX model*.json: metadata needed by the runtimeVariant directories:
AMI/CALLHOME/DIHARD II/DIHARD III/| Variant | Package | Configured max speakers | Model output capacity |
|---|---|---|---|
| AMI | AMI/ls_eend_ami_step.onnx | 4 | 6 |
| CALLHOME | CALLHOME/ls_eend_callhome_step.onnx | 7 | 9 |
| DIHARD II | DIHARD II/ls_eend_dih2_step.onnx | 10 | 12 |
| DIHARD III | DIHARD III/ls_eend_dih3_step.onnx | 10 | 12 |
The metadata JSON distinguishes between:
max_speakers: the dataset/config speaker setting from the LS-EEND infer YAMLmax_nspks: the exported model's full decode/output capacityAll four non-quantized exports in this repo use the same frontend settings:
8000 Hz200 samples80 samples102423710logmel23_cummn10 Hzfloat32These are step-wise streaming models. A runtime must maintain and feed the recurrent state tensors between calls:
enc_ret_kvenc_ret_scaleenc_conv_cachedec_ret_kvdec_ret_scaletop_bufferThe ONNX inputs and outputs follow the LS-EEND step export used by the reference Python and Swift runtimes.
Use these packages with a runtime that:
8 kHzingest/decode control inputs to handle the encoder delay and final tail flushThis repository is not a drop-in replacement for generic Hugging Face transformers inference. It is meant for custom ONNX runtimes, such as:
Setup the virtual environment:
# Create and activate virtual environment
python3.10 -m venv venv
source venv/bin/activate
# Install dependencies
pip install -r example/requirements.txt
To run the Python microphone inference script for the DIHARD III variant, run the following command:
python example/ls_eend_onnx_mic_gui.py --onnx-model DIHARD\ III/ls_eend_dih3_step.onnx
For the other variants, replace DIHARD\ III/ls_eend_dih3_step.onnx with the path to the desired variant:
# AMI variant
python example/ls_eend_onnx_mic_gui.py --onnx-model AMI/ls_eend_ami_step.onnx
# CALLHOME variant
python example/ls_eend_onnx_mic_gui.py --onnx-model CALLHOME/ls_eend_callhome_step.onnx
# DIHARD II variant
python example/ls_eend_onnx_mic_gui.py --onnx-model DIHARD\ II/ls_eend_dih2_step.onnx
Each variant ships a sidecar JSON with fields like:
{
"sample_rate": 8000,
"win_length": 200,
"hop_length": 80,
"n_fft": 1024,
"n_mels": 23,
"context_recp": 7,
"subsampling": 10,
"feat_type": "logmel23_cummn",
"frame_hz": 10.0,
"max_speakers": 10,
"max_nspks": 12
}
Check the variant-specific *.json file for the exact state tensor shapes and output dimensions.
These ONNX exports were produced from the LS-EEND code in the FS-EEND repository:
The export path is based on the LS-EEND ONNX step exporter and variant batch exporter in that project.
From the source project, the reported real-world diarization error rates are:
| Dataset | DER (%) |
|---|---|
| CALLHOME | 12.11 |
| DIHARD II | 27.58 |
| DIHARD III | 19.61 |
| AMI Dev | 20.97 |
| AMI Eval | 20.76 |
These numbers come from the upstream LS-EEND project README and reflect the original training/evaluation setup, not a Hugging Face evaluation pipeline.
The upstream LS-EEND model/codebase used for these ONNX exports is MIT-licensed, and this repository is published as MIT accordingly.
The underlying evaluation and fine-tuning datasets still have their own access and usage terms:
This repository redistributes ONNX exports of the LS-EEND model variants. Dataset licensing and access requirements remain governed by the original dataset providers.
If you use LS-EEND, cite the original paper:
@ARTICLE{11122273,
author={Liang, Di and Li, Xiaofei},
journal={IEEE Transactions on Audio, Speech and Language Processing},
title={LS-EEND: Long-Form Streaming End-to-End Neural Diarization With Online Attractor Extraction},
year={2025},
volume={33},
pages={3568-3581},
doi={10.1109/TASLPRO.2025.3597446}
}