Downloads · 30 days
58
1% of all-time downloads
nectec/Pathumma-whisper-th-medium
Pathumma-whisper-th-medium is a automatic speech recognition model from nectec. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
58
1% of all-time downloads
All-time downloads
6.6K
Public
Parameters
764M
1.5 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors1.5 GB · 100%
From the Hugging Face model README
Additional information is needed
You can transcribe audio files using the pipeline class with the following code snippet:
import torch
from transformers import pipeline
device = "cuda" if torch.cuda.is_available() else "cpu"
torch_dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32
lang = "th"
task = "transcribe"
pipe = pipeline(
task="automatic-speech-recognition",
model="nectec/Pathumma-whisper-th-medium",
torch_dtype=torch_dtype,
device=device,
)
pipe.model.config.forced_decoder_ids = pipe.tokenizer.get_decoder_prompt_ids(language=lang, task=task)
text = pipe("audio_path.wav")["text"]
print(text)
Additional information is needed
We extend our appreciation to the research teams engaged in the creation of the open speech model, including AIResearch, BiodatLab, Looloo Technology, SCB 10X, and OpenAI. We would like to express our gratitude to Dr. Titipat Achakulwisut of BiodatLab for the evaluation pipeline. We express our gratitude to ThaiSC, or NSTDA Supercomputer Centre, for supplying the LANTA used for model training, fine-tuning, and evaluation.
Pattara Tipaksorn, Wayupuk Sommuang, Kwanchiva Thangthai
@misc{tipaksorn2024PathummaWhisper,
title = { {Pathumma Whisper Medium (TH)} },
author = { Pattara Tipaksorn and Wayupuk Sommuang and Kwanchiva Thangthai },
url = { https://huggingface.co/nectec/Pathumma-whisper-th-medium },
publisher = { Hugging Face },
year = { 2024 },
}