Downloads · 30 days
2
18% of all-time downloads
aoiandroid/mms-lid-1024-coreml-8bit
mms-lid-1024-coreml-8bit is a machine learning model from aoiandroid. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-4.0.
Core ML conversion of facebook/mms-lid-1024 for on-device speech language identification. This variant uses 8-bit palettization (k-means): smaller and more ANE-friendly than the float16 base, with minimal accuracy los…
Downloads · 30 days
2
18% of all-time downloads
All-time downloads
11
Public
Repo size
969 MB
Likes
0
Public
Click a slice to open those files.
.bin968 MB · 100%
From the Hugging Face model README
Core ML conversion of facebook/mms-lid-1024 for on-device speech language identification. This variant uses 8-bit palettization (k-means): smaller and more ANE-friendly than the float16 base, with minimal accuracy loss in practice.
(1, 160000) float32(1, 1024); argmax → class index. Map to ISO 639-3 via labels.json or mms_lid_id2label.json| File | Description |
|---|---|
mms_lid_8bit.mlpackage | Core ML model (8-bit palettized, ANE-friendly) |
labels.json | Ordered list of 1024 ISO 639-3 language codes |
mms_lid_id2label.json | Index → language code mapping |
Same as the base model: load the .mlpackage, feed 10 s of 16 kHz mono as input_values, take argmax of logits, and look up the language in labels.json. Pad/trim and chunking recommendations apply as in the base README.
Same as base: fixed 10 s input, L2 accent misclassification, English ↔ Hawaiian/Maori confusion. Use chunking and confidence threshold where appropriate.
<!-- BEGIN_MMS_LID_MAC_TEST -->On-device smoke run: each file under INPUT/audio was resampled to 16 kHz mono float32, padded or trimmed to 160,000 samples (10 s), then passed to input_values; pred is ISO 639-3 from argmax(logits); conf is softmax mass on the predicted class (runner-side).
Note: Filenames are hints only (e.g. English.mp3 is not ground truth). Low conf or known MMS-LID confusions (e.g. English vs haw) may still appear.
MMS-LID 1024 Core ML 8-bit — Mac smoke test
Model: https://huggingface.co/aoiandroid/mms-lid-1024-coreml-8bit
Model dir: $PROJECT_ROOT/Log/mms_lid_1024_8bit_mac_test/model_repo
Audio dir: $PROJECT_ROOT/INPUT/audio
Compiled temp: /var/folders/ky/nmbswxzs0s79wdxndfw1y6wh0000gn/T/model_repo.mlmodelc
Compute: MLComputeUnits(rawValue: 2)
Input: input_values Output: logits
Labels: 1024
Host: ams-macbook-air.local macOS: Version 26.3.1 (a) (Build 25D771280a)
English.mp3 pcm_samples=9054841 pred=haw conf=0.2388 max_logit=7.5508 time_ms=1969.7
Euskara.mp3 pcm_samples=1865769 pred=hin conf=0.3650 max_logit=7.7617 time_ms=427.3
Guaraní.mp3 pcm_samples=1682285 pred=grn conf=0.9993 max_logit=14.7812 time_ms=419.7
Yorùbá.mp3 pcm_samples=1067049 pred=haw conf=0.4466 max_logit=7.9453 time_ms=392.9
afrikaasns.mp3 pcm_samples=2387800 pred=nld conf=0.9994 max_logit=15.0156 time_ms=453.2
arabic.mp3 pcm_samples=2060120 pred=ara conf=0.9989 max_logit=14.3047 time_ms=436.0
bengali.m4a pcm_samples=7836432 pred=ben conf=0.9985 max_logit=14.5000 time_ms=629.3
chinese.mp3 pcm_samples=12904245 pred=cmn conf=0.9993 max_logit=14.3438 time_ms=1184.4
isiZulu.mp3 pcm_samples=1396819 pred=heb conf=0.5809 max_logit=8.2422 time_ms=403.6
kiswahili.mp3 pcm_samples=1888757 pred=swh conf=0.9989 max_logit=14.2891 time_ms=442.2
korean.mp3 pcm_samples=2364395 pred=kor conf=0.9995 max_logit=15.2031 time_ms=477.8
russinan.m4a pcm_samples=15431029 pred=rus conf=0.2780 max_logit=7.8711 time_ms=768.5
test.mp3 pcm_samples=274560 pred=jpn conf=0.9984 max_logit=14.5391 time_ms=391.9
日本語.mp3 pcm_samples=1798234 pred=jpn conf=0.9984 max_logit=14.5625 time_ms=471.3
</details>
<!-- END_MMS_LID_MAC_TEST -->
CC-BY-NC-4.0 (inherited from facebook/mms-lid-1024).
@article{pratap2023mms,
title={Scaling Speech Technology to 1,000+ Languages},
author={Pratap, Vineel and others},
journal={arXiv preprint arXiv:2305.13516},
year={2023}
}