Downloads · 30 days
344
2% of all-time downloads
utter-project/mHuBERT-147-base-2nd-iter
mHuBERT-147-base-2nd-iter is a feature extraction model from utter-project. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as cc-by-nc-sa-4.0.
This repository contains the SECOND ITERATION mHuBERT-147 model. The best mHuBERT-147 model is available here.
Downloads · 30 days
344
2% of all-time downloads
All-time downloads
21.8K
Public
Parameters
94.4M
6.6 GB on disk
Likes
3
Public
Click a slice to open those files.
.index4.7 GB · 71%
From the Hugging Face model README
This repository contains the SECOND ITERATION mHuBERT-147 model. The best mHuBERT-147 model is available here.
MODEL DETAILS: 2nd iteration, K=1000, HuBERT base architecture (95M parameters), 147 languages.
mHuBERT-147 are compact and competitive multilingual HuBERT models trained on 90K hours of open-license data in 147 languages. Different from traditional HuBERTs, mHuBERT-147 models are trained using faiss IVF discrete speech units. Training employs a two-level language, data source up-sampling during training. See more information in our paper.
This repository contains:
Related Models:
Manifest list available here. Please note that since training, there were CommonVoice removal requests. This means that some of the listed files are no longer available.
Fairseq fork contains the scripts for training with multilingual batching with two-level up-sampling.
mHubert-147 reaches second and first position in the 10min and 1h leaderboards respectively. We achieve new SOTA scores for three LID tasks. See more information in our paper.

Datasets: For ASR/ST/TTS datasets, only train set is used.
Languages present not indexed by Huggingface: Asturian (ast), Basaa (bas), Cebuano (ceb), Central Kurdish/Sorani (ckb), Hakha Chin (cnh), Hawaiian (haw), Upper Sorbian (hsb) Kabyle (kab), Moksha (mdf), Meadow Mari (mhr), Hill Mari (mrj), Erzya (myv), Taiwanese Hokkien (nan-tw), Sursilvan (rm-sursilv), Vallader (rm-vallader), Sakha (sah), Santali (sat), Scots (sco), Saraiki (skr), Tigre (tig), Tok Pisin (tpi), Akwapen Twi (tw-akuapem), Asante Twi (tw-asante), Votic (vot), Waray (war), Cantonese (yue).
@inproceedings{boito2024mhubert,
author={Marcely Zanon Boito, Vivek Iyer, Nikolaos Lagos, Laurent Besacier, Ioan Calapodescu},
title={{mHuBERT-147: A Compact Multilingual HuBERT Model}},
year=2024,
booktitle={Interspeech 2024},
}
<img src="https://cdn-uploads.huggingface.co/production/uploads/62262e19d36494a6f743a28d/HbzC1C-uHe25ewTy2wyoK.png" width=7% height=7%>
This is an output of the European Project UTTER (Unified Transcription and Translation for Extended Reality) funded by European Union’s Horizon Europe Research and Innovation programme under grant agreement number 101070631.
For more information please visit https://he-utter.eu/
NAVER LABS Europe: https://europe.naverlabs.com/