Downloads · 30 days
7
30% of all-time downloads
retrina0678/miimo-efficientat-target4
miimo-efficientat-target4 is a audio classification model from retrina0678. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as other.
EfficientAT mn10as 를 4개 음향 이벤트로 파인튜닝한 모델. 2초 클립 단위 분류이고 엣지 디바이스(라즈베리파이) 배포를 염두에 뒀다.
Downloads · 30 days
7
30% of all-time downloads
All-time downloads
23
Public
Repo size
85.1 MB
Likes
0
Public
Click a slice to open those files.
.pt102 MB · 95%
From the Hugging Face model README
EfficientAT mn10_as 를 4개 음향 이벤트로
파인튜닝한 모델. 2초 클립 단위 분류이고 엣지 디바이스(라즈베리파이) 배포를 염두에 뒀다.
baby_cry, bicycle, glass_break, gunshotmn10_as (MobileNetV3, AudioSet 사전학습) + MLP head| metric | mean | std | min | max |
|---|---|---|---|---|
| accuracy | 0.9774 | 0.0113 | 0.9617 | 0.9909 |
| balanced accuracy | 0.9763 | 0.0113 | 0.9618 | 0.9903 |
| macro precision | 0.9758 | 0.0123 | 0.9594 | 0.9906 |
| macro recall | 0.9763 | 0.0113 | 0.9618 | 0.9903 |
| macro F1 | 0.9756 | 0.0122 | 0.9592 | 0.9904 |
| 클래스 | mean | std | 비고 |
|---|---|---|---|
bicycle | 0.9980 | 0.0046 | |
baby_cry | 0.9861 | 0.0113 | |
gunshot | 0.9849 | 0.0070 | |
glass_break | 0.9362 | 0.0314 | ⚠️ 가장 약함 — 주로 gunshot 으로 오분류 |
전체 5-fold pooled confusion matrix에서 glass_break → gunshot 오분류가 36건으로
가장 큰 오차 원인이다. 파열음 계열이라 2초 창에서 혼동되는 것으로 보인다.
| 클래스 | 출처 | recall | n |
|---|---|---|---|
| baby_cry | AI Hub (증강) | 0.9744 | 585 |
| baby_cry | ESC-50 (증강) | 1.0000 | 305 |
| baby_cry | donateacry | 1.0000 | 190 |
| bicycle | AI Hub (증강) | 0.9980 | 490 |
| glass_break | AI Hub (증강) | 0.9362 | 580 |
| gunshot | AI Hub (증강) | 0.9849 | 595 |
best_model.pt 최종 체크포인트 (fold 2, best by macro_recall)
config.json 학습 하이퍼파라미터
labels.json 클래스 순서 / label2id
all_metrics.json fold별 전체 지표
kfold_summary.csv fold별 요약
recall_by_source.csv 출처별 recall
report.md 상세 리포트 (혼동행렬 포함)
fold_01..05/ fold별 best.pt + 예측·혼동행렬
best_model.pt 는 dict이고 키는 다음과 같다:
model_state_dict, config, labels, label2id, id2label, fold, epoch, val_metrics
EfficientAT 저장소가 필요하다.
git clone https://github.com/fschmid56/EfficientAT
pip install torch torchaudio huggingface_hub
import sys, torch, torchaudio
from huggingface_hub import hf_hub_download
sys.path.insert(0, "EfficientAT")
from models.mn.model import get_model as get_mn
from models.preprocess import AugmentMelSTFT
from helpers.utils import NAME_TO_WIDTH
ckpt_path = hf_hub_download("retrina0678/miimo-efficientat-target4", "best_model.pt")
ckpt = torch.load(ckpt_path, map_location="cpu", weights_only=False)
labels = ckpt["labels"] # ['baby_cry','bicycle','glass_break','gunshot']
model = get_mn(num_classes=len(labels), pretrained_name="mn10_as",
width_mult=NAME_TO_WIDTH("mn10_as"), head_type="mlp")
model.load_state_dict(ckpt["model_state_dict"])
model.eval()
# 학습과 동일한 전처리 — freqm/timem 은 추론 시 0 으로 둔다
mel = AugmentMelSTFT(n_mels=128, sr=32000, win_length=800, hopsize=320, freqm=0, timem=0)
mel.eval()
wav, sr = torchaudio.load("clip.wav")
if sr != 32000:
wav = torchaudio.functional.resample(wav, sr, 32000)
wav = wav.mean(0, keepdim=True)[:, :64000] # mono, 2초
with torch.no_grad():
logits = model(mel(wav).unsqueeze(1))
if isinstance(logits, (tuple, list)):
logits = logits[0]
probs = logits.flatten(1).softmax(-1)[0]
for lbl, p in sorted(zip(labels, probs.tolist()), key=lambda x: -x[1]):
print(f"{lbl:<12} {p:.4f}")
| 항목 | 값 |
|---|---|
| 백본 | mn10_as (AudioSet 사전학습) |
| head | mlp |
| epochs / batch | 20 / 16 |
| optimizer | lr 1e-4, weight decay 1e-4 |
| freeze backbone | 앞 2 epoch |
| CV | StratifiedGroupKFold 5-fold, test_size 0.2 |
| group column | group_id (= aug_<source_file>) |
| selection metric | macro_recall |
| class weight | balanced + weighted sampler |
| seed | 42 |
| AMP | on |
데이터 누수 방지: 같은 원본 음원에서 나온 클립은 50% 오버랩 + 증강으로 서로 겹치므로,
클립이 아니라 source_file 단위(group_id)로 분할했다.
최소 클래스(bicycle 3,747)에 맞춰 클래스당 3,747개로 downsample → 총 14,988 클립.
glass_break recall이 0.936으로 가장 낮고 fold 간 편차(std 0.031)도 크다. gunshot 과의 혼동이 주 원인.recall_by_source.csv 참고.⚠️ 상업적 이용 불가.
학습 데이터에 ESC-50(CC BY-NC 3.0)과 AI Hub 제공 데이터(재배포 제한)가 포함되어 있다. 이 가중치는 해당 데이터에서 파생되었으므로 원 데이터의 제약을 그대로 승계한다.
mn10_as 자체의 라이선스는 EfficientAT 저장소를 따른다.EfficientAT: F. Schmid, K. Koutini, G. Widmer.
"Efficient Large-Scale Audio Tagging via Transformer-to-CNN Knowledge Distillation." ICASSP 2023.
ESC-50: K. J. Piczak. "ESC: Dataset for Environmental Sound Classification." ACM MM 2015.