Downloads · 30 days
44
30% of all-time downloads
Jvachier/dcase2020-task2-anomaly-detection
dcase2020-task2-anomaly-detection is a machine learning model from Jvachier. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for keras. The card lists the license as apache-2.0.
Unsupervised acoustic anomaly detection for industrial valves. Trained on normal operating sounds only (no anomalous examples seen during training); anomalies are detected as embeddings that fall far from the learned…
Downloads · 30 days
44
30% of all-time downloads
All-time downloads
149
Public
Repo size
6.3 MB
Likes
0
Public
Click a slice to open those files.
.keras5.2 MB · 83%
From the Hugging Face model README
Unsupervised acoustic anomaly detection for industrial valves. Trained on normal operating sounds only (no anomalous examples seen during training); anomalies are detected as embeddings that fall far from the learned distribution of normal acoustic signatures.
The model is a multiscale Transformer encoder trained on a pretext task: given a
log-mel spectrogram patch, predict which physical valve unit (machine ID) it came from.
This forces the encoder to learn fine-grained acoustic signatures characteristic of each
unit's normal operation. An ArcFace additive angular margin loss (s=32, m=0.10)
tightens the per-ID clusters on the embedding hypersphere during training.
After training, the classification head is discarded. What remains is a 128-dimensional embedding function. At inference, the anomaly score is the Mahalanobis distance from a clip's embedding to a single LedoitWolf-estimated Gaussian fitted over all training embeddings pooled across machine IDs — this is what makes the deployed model usable without knowing which physical unit produced a given recording.
Architecture:
Input: log-mel spectrogram (64 frames x 128 mel bins)
-> multiscale patch tokenisation: patch sizes [8, 4] frames, each scale gets its own
linear projection to 128-d + sinusoidal positional encoding, then concatenated
-> 2 x Transformer encoder block (4 heads, d_model=128)
-> mean pooling -> 128-d embedding
-> (training only) ArcFace head, s=32, m=0.10, over 4 machine-ID classes
Trained independently for 3 random seeds (0, 1, 2); the reference implementation z-score-ensembles all three at inference for a more robust score (see Usage below).
This model uses custom Keras layers (ArcFaceLayer, multiscale patch tokenizer) that
are registered via @keras.saving.register_keras_serializable in the companion
iaad Python package — loading the .keras file with bare keras.models.load_model()
without that package installed will fail to resolve the custom layers. Install the
package from GitHub (not yet published to PyPI):
pip install git+https://github.com/jvachier/industrial-audio-anomaly-detection.git
Single seed:
from iaad.inference import AnomalyDetector
det = AnomalyDetector.from_hub(machine_type="valve", seed=0)
result = det.score_file("path/to/valve_clip.wav", score="mahalanobis")
print(result.is_anomalous, result.percentile)
Ensemble of all 3 seeds (recommended — more robust score):
from iaad.inference import EnsembleDetector
ens = EnsembleDetector.from_hub(machine_type="valve")
result = ens.score_file("path/to/valve_clip.wav", score="mahalanobis")
print(result["is_anomalous"], result["percentile"])
Input audio should be a 16 kHz mono WAV clip (10 s recommended, matching training clip length); other sample rates are auto-resampled with a warning.
DCASE 2020 Challenge Task 2 (MIMII-based), valve subset, mirrored on Kaggle: daisukelab/dc2020task2.
id_00, id_02, id_04, id_06Preprocessing (src/iaad/features.py, no external audio library):
log-mel spectrogram, 128 mel bins, 50–8000 Hz, 512-sample FFT window / 256-sample hop,
64-frame (~1.02 s) non-overlapping patches, per-clip CMVN (mean-normalised per mel bin).
SpecAugment is disabled for valve specifically (empirically best AUC without it; fan/pump
use frequency + time masking).
| Hyperparameter | Value |
|---|---|
| Optimizer | Adam, lr=1e-3 |
| Batch size | 64 |
| Epochs | up to 120, early stopping (patience=10, restore best weights) |
| Loss | ArcFace-modified softmax cross-entropy over 4 machine-ID classes |
| ArcFace scale / margin | s=32, m=0.10 |
| Seeds | 0, 1, 2 (independent runs, deterministic — tf.config.experimental.enable_op_determinism()) |
| Hardware | single GPU (Kaggle P100/T4), no mixed precision |
After training, a calibration pass fits a LedoitWolf covariance estimator over all
training-set embeddings (pooled across machine IDs) and records percentile thresholds
of the resulting Mahalanobis distances — this is what ships in valve_calib_s{seed}.npz
and lets inference flag a clip as anomalous (e.g. above the 95th percentile of normal)
without any refitting.
Per-machine-ID mean AUC-ROC, averaged over 3 random seeds (chance = 0.500):
| Machine | EMB pooled Mahalanobis |
|---|---|
| valve | 0.947 ± 0.006 |
Pooled scoring is the deployment-safe metric (no machine ID needed at inference). See the GitHub repo's evaluation docs for the full comparison against earlier detector families (dense autoencoder, Transformer-VAE) and the oracle (per-ID, ID-known) scoring variant.
| File | Description |
|---|---|
version.json | {"emb_version": "v5", "run_tag": "v5.4"} — model family marker read by AnomalyDetector._load() |
valve_emb_s{0,1,2}.keras | Trained EMBClassifier weights, one per seed |
valve_norm_s{0,1,2}.npz | Per-mel-bin normalisation stats (mean, std) fit on that seed's training split |
valve_calib_s{0,1,2}.npz | Calibration: pooled LedoitWolf location/precision + percentile thresholds |
Apache 2.0 — see the GitHub repository's LICENSE.
@misc{vachier2025iaad,
author = {Vachier, Jeremy},
title = {Industrial Audio Anomaly Detection -- DCASE 2020 Task 2 Baseline},
year = {2026},
url = {https://github.com/jvachier/industrial-audio-anomaly-detection},
}
@inproceedings{koizumi2020dcase,
author = {Koizumi, Yuma and Saito, Shoichiro and Uematsu, Hisashi and Harada, Noboru and Imoto, Keisuke},
title = {ToyADMOS: A Dataset of Miniature-Machine Operating Sounds for Anomalous Sound Detection},
booktitle = {Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020)},
year = {2020},
}
@misc{huggingface_hub,
author = {Hugging Face},
title = {Hugging Face Hub},
year = {2026},
url = {https://github.com/huggingface/huggingface_hub},
}