Downloads · 30 days
0
drakulavich/SpeechBrain-coreml
SpeechBrain-coreml is a audio classification model from drakulavich. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for onnx. The card lists the license as apache-2.0.
Pre-converted speechbrain/lang-id-voxlingua107-ecapa model for spoken language identification. Supports 107 languages from audio.
Downloads · 30 days
0
Access
Public
Updated Apr 14, 2026
Repo size
126 MB
Likes
0
Public
Click a slice to open those files.
.data85.3 MB · 68%
From the Hugging Face model README
Pre-converted speechbrain/lang-id-voxlingua107-ecapa model for spoken language identification. Supports 107 languages from audio.
Converted for use with Kesha Voice Kit — open-source voice toolkit.
| File | Format | Size | Description |
|---|---|---|---|
lang-id-ecapa.onnx | ONNX | ~760KB | Model graph |
lang-id-ecapa.onnx.data | ONNX | ~85MB | Model weights (external data) |
lang-id-ecapa.mlpackage.tar.gz | CoreML | ~40MB | CoreML model archive (macOS) |
labels.json | JSON | <1KB | 107 ISO 639-1 language codes |
bun install -g @drakulavich/kesha-voice-kit
kesha install # downloads this model automatically
kesha --json audio.ogg # transcribe + detect language
import onnxruntime as ort
import numpy as np
import json
session = ort.InferenceSession("lang-id-ecapa.onnx")
with open("labels.json") as f:
labels = json.load(f)
# Input: 16kHz mono float32 waveform
audio = np.random.randn(1, 160000).astype(np.float32) # 10 seconds
result = session.run(None, {"waveform": audio})
probs = result[0][0]
top_idx = np.argmax(probs)
print(f"Language: {labels[top_idx]} (confidence: {probs[top_idx]:.4f})")
use ort::session::Session;
let session = Session::builder()?.commit_from_file("lang-id-ecapa.onnx")?;
// Input: "waveform" [1, samples] float32
// Output: "language_probs" [1, 107] float32
[1, samples] float32)[1, 107] float32, softmax applied)ab, af, am, ar, as, az, ba, be, bg, bn, bo, br, ca, ceb, cs, cy, da, de, el, en, eo, es, et, eu, fa, fi, fo, fr, gl, gn, gu, ha, haw, he, hi, hr, ht, hu, hy, ia, id, is, it, ja, jw, ka, kk, km, kn, ko, la, lb, ln, lo, lt, lv, mg, mi, mk, ml, mn, mr, ms, mt, my, ne, nl, nn, no, oc, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, sn, so, sq, sr, su, sv, sw, ta, te, tg, th, tk, tl, tr, tt, uk, ur, uz, vi, war, yi, yo, zh
Converted from PyTorch using torch.onnx.export (ONNX) and torch.export + coremltools (CoreML).
Conversion script: scripts/convert-lang-id-model.py
Apache 2.0 (same as the original SpeechBrain model)