Downloads · 30 days
638
88% of all-time downloads
eliya/forensics_0.3B_base_deepfake_classifier
forensics_0.3B_base_deepfake_classifier is a audio classification model from eliya. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as cc-by-nc-4.0.
The default speech deepfake detector of the Forensics family. WavLM-large + AASIST graph-attention, fully fine-tuned end-to-end (no frozen shortcuts) across a wide multi-source mix of TTS spoofs, voice conversion, cod…
Downloads · 30 days
638
88% of all-time downloads
All-time downloads
724
Public
Repo size
5.1 GB
Likes
7
Public
Click a slice to open those files.
.pt3.8 GB · 75%
From the Hugging Face model README
The default speech deepfake detector of the Forensics family. WavLM-large + AASIST graph-attention, fully fine-tuned end-to-end (no frozen shortcuts) across a wide multi-source mix of TTS spoofs, voice conversion, codec artifacts, and the standard anti-spoofing benchmark suite. A combined cross-entropy + OC-Softmax + supervised-contrastive objective gives it a decision boundary that holds up well outside its own training distribution — not just on the data it saw.
Feed it 5 seconds of audio, get back a calibrated real/fake probability. Sub-1% EER on held-out data, and under 2% across most of a 14-benchmark external sweep (ASVspoof, ADD, In-the-Wild, LibriSeVoc, SONAR, and more).
microsoft/wavlm-large (~300M params, fully unfrozen)| Model | Use it for |
|---|---|
forensics_0.3B_base_deepfake_classifier (this model) | general-purpose default |
forensics_0.3B_xlsr_wild_deepfake_classifier | uncontrolled / real-world audio |
forensics_0.3B_v2_deepfake_age_gender_classifier 🆕 | speaker age/gender, hardened against the newest TTS threats — our latest release |
forensics_0.3B_wavlm_oc_softmax_deepfake_classifier | tighter bonafide boundary, ensembling |
Full family: huggingface.co/collections/eliya/forensics-speech-deepfake-detection-family
Trained using an agentic training loop — see eliyasegev/autotrain.
Fine-tuned with a combined loss for robustness beyond any single objective:
Trained on 5M+ speech samples spanning a mix of proprietary and publicly available data — some of our own self-created data has been open-sourced as well. Built for robustness against a wide range of TTS generation methods — including some of the most advanced synthesis techniques available today. 5-second crops, AdamW, cosine LR schedule, and a heavy augmentation stack — codec transcoding (mp3/aac/opus/vorbis/µ-law/A-law/GSM), noise augmentation, RIR, RawBoost, SpecAugment, FreqMask, splice/mix, and cross-class splice — so the model sees more distortion during training than it will ever encounter in the wild.
| Eval set | EER % |
|---|---|
| Val (held-out) | 0.72 |
| MLAAD (v7) | 0.71 |
| CodecFake | 0.54 |
| DFADD | 0.00 |
| MD-CommonVoice | 0.17 |
| In-the-Wild | 1.38 |
| ASVspoof2019-LA | 0.26 |
| ASVspoof2021-LA | 1.56 |
| ASVspoof2024 | 11.91 |
| ADD2022-Track1 | 17.34 |
| ADD2022-Track3 | 3.03 |
| ADD2023-Round1 | 6.46 |
| ADD2023-Round2 | 13.00 |
| LibriSeVoc | 0.04 |
| SONAR | 0.44 |
| Avg (all sets) | 3.84 |
| Avg (external only) | 4.06 |
Consistently sub-2% EER across almost every external benchmark, with strong results even on the harder ADD/ASVspoof2024 tracks.
| file | purpose |
|---|---|
checkpoint_epoch_5.safetensors | model weights, safe format |
checkpoint_epoch_5.pt | model weights, legacy pickle |
config.json | minimal architecture metadata (also used by the Hub to track downloads) |
inference.py | run script — prefers the .safetensors file automatically |
model.py | architecture |
requirements.txt | deps |
pip install -r requirements.txt # torch, torchaudio, transformers, safetensors
hf download eliya/forensics_0.3B_base_deepfake_classifier --local-dir .
python inference.py <audio.wav>
(Optionally override the checkpoint: python inference.py <audio.wav> <checkpoint.pt>.)
Audio is auto-converted to mono / 16 kHz and trimmed/padded to 5 s.
fake_probability: <0..1> # threshold is domain-dependent — adjust to your use case; ~0.1-0.2 is usually the best range
bonafide_score: <0..1> # raw P(real)
verdict: REAL | FAKE
$ python inference.py real_human.wav
fake_probability: 0.0503
bonafide_score: 0.9497
verdict: REAL
$ python inference.py tts_fake.wav
fake_probability: 0.8641
bonafide_score: 0.1359
verdict: FAKE
Higher fake_probability = more likely a deepfake. Score is 1 − sigmoid(logit),
since the classifier is trained with label 1 = real, 0 = fake.
CC-BY-NC-4.0 — free for personal and research use.
If you use this in research, evaluation, or training, please acknowledge it:
@misc{segev2026forensicsbase,
author = {Segev, Eliya},
title = {Forensics 0.3B: Speech Deepfake Detection Classifier (Base)},
year = {2026}
}