Downloads · 30 days
76
5% of all-time downloads
bioamla/ast-esc50
ast-esc50 is a audio classification model from bioamla. Use it for the audio classification task on the model card, and read the license before you ship it in a product. The card lists the license as bsd-3-clause.
An Audio Spectrogram Transformer (AST) model fine-tuned on the ESC-50 dataset for environmental sound classification.
Downloads · 30 days
76
5% of all-time downloads
All-time downloads
1.7K
Public
Parameters
86.2M
345 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors345 MB · 100%
From the Hugging Face model README
An Audio Spectrogram Transformer (AST) model fine-tuned on the ESC-50 dataset for environmental sound classification.
This model is based on the Audio Spectrogram Transformer architecture, fine-tuned to classify 50 categories of environmental sounds. The AST applies a pure attention mechanism to audio spectrograms, treating them as sequences of patches similar to Vision Transformers (ViT).
The model classifies audio into 50 environmental sound categories:
Animals: cat, chirping_birds, cow, crow, dog, frog, hen, insects, pig, rooster, sheep
Natural Sounds: crackling_fire, crickets, rain, sea_waves, thunderstorm, water_drops, wind
Human Sounds: breathing, brushing_teeth, clapping, coughing, crying_baby, drinking_sipping, footsteps, laughing, sneezing, snoring
Domestic Sounds: clock_alarm, clock_tick, door_wood_creaks, door_wood_knock, glass_breaking, keyboard_typing, mouse_click, toilet_flush, vacuum_cleaner, washing_machine
Urban Sounds: airplane, car_horn, church_bells, engine, fireworks, helicopter, siren, train
Mechanical/Tools: can_opening, chainsaw, hand_saw, pouring_water
BSD-3-Clause