Downloads · 30 days
185
2% of all-time downloads
biodatlab/whisper-th-medium-timestamp
whisper-th-medium-timestamp is a automatic speech recognition model from biodatlab. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of biodatlab/whisper-th-medium-combined on a custom-created longform dataset derived from the CMKL/Porjai-Thai-voice-dataset-central. It achieves the following results on the common-…
Downloads · 30 days
185
2% of all-time downloads
All-time downloads
8K
Public
Parameters
764M
3.1 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors3.1 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of biodatlab/whisper-th-medium-combined on a custom-created longform dataset derived from the CMKL/Porjai-Thai-voice-dataset-central. It achieves the following results on the common-voice-13 test set:
This model is designed to perform automatic speech recognition (ASR) for the Thai language, with the added capability of generating timestamps for the transcribed text. It's based on the Whisper medium architecture and has been fine-tuned on a specially crafted dataset to enable timestamp generation.
Use the model with Hugging Face's transformers as follows:
from transformers import pipeline
import torch
MODEL_NAME = "biodatlab/whisper-th-medium-timestamp" # specify the model name
lang = "th" # Thai language
device = 0 if torch.cuda.is_available() else "cpu"
pipe = pipeline(
task="automatic-speech-recognition",
model=MODEL_NAME,
chunk_length_s=30,
device=device,
return_timestamps=True,
)
pipe.model.config.forced_decoder_ids = pipe.tokenizer.get_decoder_prompt_ids(
language=lang,
task="transcribe"
)
result = pipe("audio.mp3", return_timestamps=True)
text = result["text"]
timestamps = result["chunks"]
This model is intended for Thai automatic speech recognition tasks, particularly where timestamp information is required. It can be used for transcribing Thai audio content, creating subtitles, or any application that needs to align text with specific time points in audio. The model's performance on speech recognition may be lower compared to non-timestamped versions due to the additional complexity of the task and the pseudo-timestamp generation method used in training.
The model was trained on a custom-created longform dataset derived from the CMKL/Porjai-Thai-voice-dataset-central. The dataset creation process involved the following steps:
This approach allowed us to create a dataset with longer, more diverse audio samples and approximate timestamp information, which is crucial for training a model capable of generating timestamps.
The model was fine-tuned using a custom training script that incorporates the following:
The following hyperparameters were used during training:
The WER (Word Error Rate) of 15.57 on the Common Voice 13 test set indicates good performance for Thai ASR. However, it's important to note that the timestamp generation model has a lower accuracy compared to the non-timestamped version of the model. This is due to several factors:
Users should be aware that while the timestamps provide a general indication of when words or phrases occur in the audio, they may not be as precise as manually annotated timestamps. The model's performance may also vary depending on the acoustic conditions, speaker variability, and the presence of background noise in the input audio.
If you use this model in your research or applications, please cite it as follows:
@misc{biodatlab_whisper_th_medium_timestamp,
author = {Atirut Boribalburephan, Zaw Htet Aung, Knot Pipatsrisawat, Titipat Achakulvisut},
title = {Whisper Medium Thai Timestamp: A fine-tuned Whisper model for Thai automatic speech recognition with timestamp generation},
year = 2024,
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/biodatlab/whisper-th-medium-timestamp}}
}