Downloads Β· 30 days
0
rohitkumarai/deepfake_audio_classifier
deepfake_audio_classifier is a machine learning model from rohitkumarai. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A machine learning model to detect deepfake/synthetic audio using Wav2Vec2 embeddings and classical ML classifiers.
Downloads Β· 30 days
0
Access
Public
Updated Dec 7, 2025
Repo size
50.7 KB
Likes
0
Public
Click a slice to open those files.
.pkl50.7 KB Β· 85%
From the Hugging Face model README
A machine learning model to detect deepfake/synthetic audio using Wav2Vec2 embeddings and classical ML classifiers.
| Model | Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|---|
| Logistic Regression | 92.86% | 0.95 | 0.93 | 0.93 |
| SVM | 85.71% | 0.89 | 0.86 | 0.85 |
| Random Forest | 78.57% | 0.85 | 0.79 | 0.76 |
Best Model: Logistic Regression with 92.86% accuracy
We use Wav2Vec2 (facebook/wav2vec2-base-960h) to extract deep audio embeddings:
Pipeline:
Audio File β Wav2Vec2 β 768-dim Embedding β Classifier β Prediction
Three classifiers were trained and compared:
pip install transformers torch librosa soundfile scikit-learn huggingface-hub requests numpy
from predict_from_hf import AudioDeepfakeDetectorFromHF
# Initialize detector (downloads model automatically)
detector = AudioDeepfakeDetectorFromHF("hjsgfd/deepfake_audio_classifier")
# Predict from URL
result = detector.predict("https://your-audio-file.wav", is_url=True)
print(f"Prediction: {result['label']} ({result['confidence']:.1%})")
from predict_from_hf import AudioDeepfakeDetectorFromHF
detector = AudioDeepfakeDetectorFromHF("hjsgfd/deepfake_audio_classifier")
# Multiple URLs
audio_urls = [
"https://example.com/audio1.wav",
"https://example.com/audio2.wav",
"https://example.com/audio3.wav",
]
results = detector.predict_batch(audio_urls, are_urls=True)
# Print results
for result in results:
if 'prediction' in result:
print(f"{result['audio_source']}: {result['label']} ({result['confidence']:.1%})")
# Single file
result = detector.predict("path/to/audio.wav", is_url=False)
# Multiple files
local_files = ["audio1.wav", "audio2.wav", "audio3.wav"]
results = detector.predict_batch(local_files, are_urls=False)
The model consists of three files hosted on Hugging Face:
{
"model_type": "LogisticRegression",
"accuracy": 0.9286,
"feature_extractor": "facebook/wav2vec2-base-960h",
"embedding_dim": 768,
"num_classes": 5,
"class_labels": {
"0": "class_0",
"1": "class_1",
"2": "class_2",
"3": "class_3",
"4": "class_4"
}
}
precision recall f1-score support
class_0 1.00 0.67 0.80 3
class_1 1.00 1.00 1.00 2
class_2 1.00 1.00 1.00 3
class_3 0.75 1.00 0.86 3
class_4 1.00 1.00 1.00 3
accuracy 0.93 14
macro avg 0.95 0.93 0.93 14
weighted avg 0.95 0.93 0.93 14
transformers>=4.30.0
torch>=2.0.0
librosa>=0.10.0
soundfile>=0.12.0
scikit-learn>=1.3.0
huggingface-hub>=0.16.0
requests>=2.31.0
numpy>=1.24.0
Input: Audio File (any format supported by soundfile)
β
Preprocessing (16kHz, Mono)
β
Wav2Vec2 Feature Extractor
β
768-dimensional Embedding
β
StandardScaler Normalization
β
Logistic Regression Classifier
β
Output: Class Prediction + Confidence Scores
Contributions are welcome! Please feel free to submit a Pull Request.
If you use this model, please cite:
@misc{deepfake_audio_classifier_2024,
author = {Your Name},
title = {Deepfake Audio Detection Model},
year = {2024},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/hjsgfd/deepfake_audio_classifier}}
}
For questions or feedback, please open an issue on the repository.