Downloads · 30 days
0
marianbasti/audio-language-classification
audio-language-classification is a audio classification model from marianbasti. Use it for the audio classification task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
First iteration of a lightweight (7M parameter) model for detecting language from a speech audio. Code is available at GitHub
Downloads · 30 days
0
Access
Public
Updated Oct 21, 2025
Repo size
79.4 MB
Likes
2
Trending 1
Click a slice to open those files.
.pt159 MB · 100%
From the Hugging Face model README
First iteration of a lightweight (7M parameter) model for detecting language from a speech audio. Code is available at GitHub
In short:
Supported languages (labels):
Intended use:
Out-of-scope:
Data:
Model architecture:
Training setup:
Preprocessing:
Evaluation results:
Files and checkpoints:
How to use (inference):
import torch
from models import LanguageClassifier
device = "cuda" if torch.cuda.is_available() else "cpu"
# Build and load from a directory containing best_model.pt and language_mapping.txt
model = LanguageClassifier.from_pretrained("training", device=device)
# Single-file prediction (auto resample to 24k, pad/trim to 10s)
label, prob = model.predict("example.wav", max_length_seconds=10.0, top_k=1)
print(label, prob)
# Top-3
top3 = model.predict("example.wav", top_k=3)
print(top3) # [('en', 0.62), ('de', 0.21), ('fr', 0.08)]
# If you already have a waveform tensor:
# wav: torch.Tensor [T] at 24kHz (or provide sample_rate to auto-resample)
# model.predict handles [T] or [B,T]
# label, prob = model.predict(wav, sample_rate=orig_sr, top_k=1)
Limitations and risks:
Reproducibility:
Citation and acknowledgements: