Downloads · 30 days
6.3K
18% of all-time downloads
changelinglab/PhoneticXeus
PhoneticXeus is a automatic speech recognition model from changelinglab. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as cc-by-nc-sa-4.0.
Multilingual phone recognition that turns speech into IPA phones, built on the XEUS speech encoder. Trained on 70+ languages (IPAPack++).
Downloads · 30 days
6.3K
18% of all-time downloads
All-time downloads
34.3K
Public
Parameters
575M
11.5 GB on disk
Likes
15
Public
Click a slice to open those files.
.pt2.3 GB · 50%
From the Hugging Face model README
Multilingual phone recognition that turns speech into IPA phones, built on the XEUS speech encoder. Trained on 70+ languages (IPAPack++).
pip install torch torchaudio transformers huggingface_hub safetensors soundfile numpy pyyaml typeguard
import torchaudio
from transformers import AutoModel
model = AutoModel.from_pretrained(
"changelinglab/PhoneticXeus", trust_remote_code=True
).eval()
wav, sr = torchaudio.load("audio.wav")
wav = wav.mean(0) # mono, shape (samples,)
if sr != 16000:
wav = torchaudio.functional.resample(wav, sr, 16000)
print(model.transcribe(wav, sampling_rate=16000)[0]["processed_transcript"])
# e.g. "aɪhædðætkʰjʊɹiɑsətipɪsaɪd…"
model.transcribe(...) returns a list of dicts with processed_transcript
(joined IPA) and predicted_transcript (slash-separated phones). Calling
model(input_values) returns frame-level CTC logits (batch, frames, 428)
for custom decoding.
Audio must be mono 16 kHz. The first load asks you to allow the repo's
remote code (trust_remote_code=True).
@misc{pxeus26,
title={An Empirical Recipe for Universal Phone Recognition},
author={Shikhar Bharadwaj and Chin-Jou Li and Kwanghee Choi and Eunjung Yeo and William Chen and Shinji Watanabe and David R. Mortensen},
year={2026},
eprint={2603.29042},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.29042},
}
CC-by-NC SA 4.0, please note this is an update.