Downloads · 30 days
41
2% of all-time downloads
Cnam-LMSSC/phonemizer_headset_microphone
phonemizer_headset_microphone is a automatic speech recognition model from Cnam-LMSSC. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as mit.
<p align="center" <img src="https://cdn-uploads.huggingface.co/production/uploads/65302a613ecbe51d6a6ddcec/zhB1fh-c0pjlj-Tr4Vpmr.png" style="object-fit:contain; width:280px; height:280px;" </p
Downloads · 30 days
41
2% of all-time downloads
All-time downloads
2.1K
Public
Parameters
94.4M
378 MB on disk
Likes
4
Public
Click a slice to open those files.
.safetensors378 MB · 100%
From the Hugging Face model README
headset_microphone audio of the speech_clean subset of Cnam-LMSSC/vibravox (see VibraVox paper on arXiV)As this model is specifically trained for a speech-to-phoneme task, the output is sequence of IPA-encoded words, without punctuation. If you don't read the phonetic alphabet fluently, you can use this excellent IPA reader website to convert the transcript back to audio synthetic speech in order to check the quality of the phonetic transcription.
An entry point to all phonemizers models trained on different sensor data from the Vibravox dataset is available at https://huggingface.co/Cnam-LMSSC/vibravox_phonemizers.
Each of these models has been trained for a specific non-conventional speech sensor and is intended to be used with in-domain data. The only exception is the headset microphone phonemizer, which can certainly be used for many applications using audio data captured by airborne microphones.
Please be advised that using these models outside their intended sensor data may result in suboptimal performance.
The model has been finetuned for 10 epochs with a constant learning rate of 1e-5. To reproduce experiment please visit jhauret/vibravox.
import torch, torchaudio
from transformers import AutoProcessor, AutoModelForCTC
from datasets import load_dataset
processor = AutoProcessor.from_pretrained("Cnam-LMSSC/phonemizer_headset_microphone")
model = AutoModelForCTC.from_pretrained("Cnam-LMSSC/phonemizer_headset_microphone")
test_dataset = load_dataset("Cnam-LMSSC/vibravox", "speech_clean", split="test", streaming=True)
audio_48kHz = torch.Tensor(next(iter(test_dataset))["audio.headset_microphone"]["array"])
audio_16kHz = torchaudio.functional.resample(audio_48kHz, orig_freq=48_000, new_freq=16_000)
inputs = processor(audio_16kHz, sampling_rate=16_000, return_tensors="pt")
logits = model(inputs.input_values).logits
predicted_ids = torch.argmax(logits,dim = -1)
transcription = processor.batch_decode(predicted_ids)
print("Phonetic transcription : ", transcription)