Downloads · 30 days
0
Inferencelab/ECAPA-TDNN-VHE
ECAPA-TDNN-VHE is a machine learning model from Inferencelab. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
- Model name: ECAPA-TDNN-VHE - Author: Muhammad Khubaib Ahmad et al. - License: Apache 2.0 - Framework: PyTorch, SpeechBrain - Embedding dimensionality: 192 - Sampling rate: 16 kHz (mono) - Task: Health-centric vocal…
Downloads · 30 days
0
Access
Public
Updated May 18, 2026
Repo size
9.4 MB
Likes
1
Public
Click a slice to open those files.
.pth9.2 MB · 98%
From the Hugging Face model README
ECAPA-TDNN-VHE (Vocal Health Encoder) is a research-grade deep neural speech encoder developed in the research of Muhammad Khubaib Ahmad for generating health-centric, speaker-invariant vocal embeddings. Unlike conventional speaker embedding models optimized for identity discrimination, ECAPA-TDNN-VHE is trained from scratch using supervised contrastive learning, explicitly promoting separation between vocal health states while minimizing speaker-specific information.
Empirical evaluation demonstrates that ECAPA-TDNN-VHE outperforms the baseline ECAPA-TDNN by over 2.5× in classification accuracy and F1-score on vocal health benchmarks, establishing it as a state-of-the-art model for health-oriented speech representation learning in ECAPA-TDNN based architectures.
The encoder forms the core of the Auralis MLOps framework and is accessible via the open-source Python library auralis_vfs, enabling reproducible and real-time vocal fatigue scoring for research and applied scenarios.
Key capabilities include:
auralis_vfs, enabling researchers to compute fatigue scores from audio files (.wav, .mp3, .m4a).This model represents a state-of-the-art (SOTA) approach for ECAPA-based health embeddings, outperforming conventional ECAPA-TDNN trained for speaker recognition.
The model was evaluated on vocal health classification tasks. Results highlight ECAPA-TDNN-VHE's superiority over baseline ECAPA-TDNN:
| Model | Accuracy | Macro F1 | Healthy F1 | Strained F1 | Stressed F1 |
|---|---|---|---|---|---|
| ECAPA-TDNN (SpeechBrain baseline) | 0.36 | 0.31 | 0.50 | 0.22 | 0.22 |
| ECAPA-TDNN-VHE (Khubaib et al., 2026) | 0.78 | 0.77 | 0.85 | 0.78 | 0.70 |
This demonstrates state-of-the-art health-centric embedding performance within ECAPA-based architectures.

Figure 1: Radar chart comparing baseline ECAPA-TDNN and ECAPA-TDNN-VHE across classification and embedding quality metrics.
| Rank | Model | Accuracy | Macro F1 |
|---|---|---|---|
| 1 | ECAPA-TDNN-VHE (Muhammad Khubaib Ahmad et al., 2026) | 0.78 | 0.77 |
| 2 | ECAPA-TDNN (SpeechBrain baseline) | 0.36 | 0.31 |
Leaderboard reflects performance on the vocal health dataset and serves as a research benchmark, not a universal ranking.
The model can be used via the Python library auralis_vfs:
pip install auralis_vfs
Example usage:
from auralis.scorer import score_audio, score_waveform
# Score from a waveform array
score = score_waveform(audio_array)
# Score from an audio file
score = score_audio("sample.wav")
print(f"Vocal fatigue score: {score:.2f}")
The model is also deployed in the Auralis MLOps system, providing real-time fatigue monitoring and embedding-based analyses.
If you use this model in your research, please cite:
@misc{muhammad_khubaib_ahmad_2026,
author = { Muhammad Khubaib Ahmad },
title = { ECAPA-TDNN-VHE (Revision 871292d) },
year = 2026,
url = { https://huggingface.co/Khubaib01/ECAPA-TDNN-VHE },
doi = { 10.57967/hf/7648 },
publisher = { Hugging Face }
}
The author gratefully acknowledge the participants for allowing us to use their voice in research and the author thank to the Data Manager(Faiez Ahmad) and Data collector(Muhammad Anas Tariq) for their incredible services and cooperation.