Downloads · 30 days
0
soorajsatheesan/audidex
audidex is a machine learning model from soorajsatheesan. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for keras. The card lists the license as mit.
This repository hosts the pre-trained Audidex model and scaler for audio deepfake detection — classifying speech as Real (genuine human) or Fake (synthetic or manipulated). The model is part of the Audidex system and…
Downloads · 30 days
0
Access
Public
Updated Feb 1, 2026
Repo size
103 MB
Likes
1
Public
Click a slice to open those files.
.h5103 MB · 100%
From the Hugging Face model README
This repository hosts the pre-trained Audidex model and scaler for audio deepfake detection — classifying speech as Real (genuine human) or Fake (synthetic or manipulated). The model is part of the Audidex system and uses hybrid Mel-spectrogram + glottal features with a fully connected neural network (FCNN).
.h5).The project addresses the rise of deepfake audio technology (identity theft, privacy concerns, deception, and fraudulent activities). The goal was to develop a detection model trained on a diverse dataset, capable of real-time detection and built for discriminating between real and synthesized speech. The approach uses hand-crafted temporal features (glottal) and frequency-domain features (Mel-spectrogram) jointly. The model is an FCNN chosen for its efficiency in processing combined feature vectors and for learning complex relationships between spectral and glottal attributes. The trained model and scaler are deployed in the Audidex backend (FastAPI) with the same feature extraction pipeline for consistent inference.
A unified and customized dataset was used to improve model accuracy and robustness:
Feature extraction is the most critical part of the detection pipeline, enabling the model to capture minute variations between authentic and fake audio.
n_mels=128, fmax=8000 Hz, amplitude to dB; time axis fixed to 128 frames; flattened to a vector and normalized (e.g. by max absolute value) before concatenation with glottal features.[flattened_mel_spectrogram, jitter, shimmer, hnr, f1, f2, f3]. Length is fixed to the model's input dimension via padding or truncation. Features are standardized using a pre-fitted scaler (scaler.pkl) at training and inference.The detection model is based on a deep learning approach. Several architectures (e.g. CNN, RNN) were investigated; an FCNN was chosen because it handles tabular-style combined feature vectors effectively and is well-matched for learning complex relationships between Mel-spectrogram and glottal features via densely connected layers.
| Layer / setting | Details |
|---|---|
| Input | 1D flattened feature vector combining spectral (Mel-spectrogram) and glottal (shimmer, jitter, HNR, formants) features. |
| First Dense | 512 neurons, ReLU. L2 regularization (λ=0.001), Batch Normalization, Dropout 40%. Output: 512-dimensional representation. |
| Second Dense | 256 neurons, ReLU. L2 regularization, Batch Normalization, Dropout 30%. Output: 256-dimensional representation. |
| Third Dense | 128 neurons, ReLU. L2 regularization, Batch Normalization. No Dropout at this stage to retain information for classification. Output: 128-dimensional representation. |
| Output | 2 neurons, Softmax → [P(Real), P(Fake)]. |
| Compilation | Optimizer: Adam (initial learning rate 0.001). Loss: Categorical cross-entropy. Metric: Accuracy. |
| Training | Dataset split 80% train / 10% validation / 10% test. Features standardized with pre-fitted scaler (scaler.pkl). 10 epochs, batch size 32, with a learning rate scheduler to reduce the learning rate over time for smoother convergence. |
0 = Real, 1 = Fake (argmax over these two classes).Convergence of loss during training indicated effective learning with low overfitting. Strong precision, recall, and F1-score on test data further validate the reliability of the model. The introduction of glottal features (shimmer, jitter, formants) greatly improved the detection capability for synthesized speech; Mel-spectrogram features captured spectral and temporal patterns well. Audio chunking in the backend supports real-time responsiveness without sacrificing classification performance. The model had some difficulty with state-of-the-art voice conversion methods that generate voices very close to real ones — a direction for future improvement.
| File | Description |
|---|---|
optimized_audio_deepfake_detector.h5 | Keras/TensorFlow FCNN model (load with tf.keras.models.load_model). |
scaler.pkl | Pre-fitted scaler (e.g. joblib.load) for standardizing the combined feature vector before inference. |
pip install huggingface_hub
from huggingface_hub import hf_hub_download
model_path = hf_hub_download(repo_id="soorajsatheesan/audidex", filename="optimized_audio_deepfake_detector.h5")
scaler_path = hf_hub_download(repo_id="soorajsatheesan/audidex", filename="scaler.pkl")
Inference must use the same feature extraction as in the Audidex backend (audio_processing.py): Mel-spectrogram (128 mels, 128 time steps) + glottal features (jitter, shimmer, HNR, 3 formants), then concatenate, pad/truncate to model input length, and apply the scaler.
from tensorflow.keras.models import load_model
from joblib import load
model = load_model(model_path)
scaler = load(scaler_path)
# Build feature vector using Audidex's audio_processing (same as backend)
# combined = [flattened_mel, jitter, shimmer, hnr, f1, f2, f3], then pad/truncate
# combined = scaler.transform(combined.reshape(1, -1))
# pred = model.predict(combined)
# label = "Real" if np.argmax(pred) == 0 else "Fake"
For full inference (including feature extraction), use the Audidex repository: clone it, place optimized_audio_deepfake_detector.h5 and scaler.pkl in backend/, and run the FastAPI app.