Downloads · 30 days
48
1% of all-time downloads
projecte-aina/whisper-large-v3-tiny-caesar
whisper-large-v3-tiny-caesar is a automatic speech recognition model from projecte-aina. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
<details <summaryClick to expand</summary
Downloads · 30 days
48
1% of all-time downloads
All-time downloads
4.8K
Public
Parameters
1.5B
17.3 GB on disk
Likes
2
Public
Click a slice to open those files.
.bin6.2 GB · 50%
From the Hugging Face model README
The "whisper-large-v3-tiny-caesar" is an acoustic model based on "openai/whisper-large-v3" suitable for Automatic Speech Recognition in code-switching conditions between Spanish and Catalan.
The "whisper-large-v3-tiny-caesar" is an acoustic model suitable for Automatic Speech Recognition in code-switching conditions between Spanish and Catalan. It is the result of fine-tuning the "openai/whisper-large-v3" with CAESAR-TINY, a 2-hour code-switching dataset in Spanish/Catalan.
This model can be used for Automatic Speech Recognition (ASR) in code-switching conditions between Spanish and Catalan. The model is intended to transcribe audio files to plain text.
To see an updated and functional version of this code, please check our Notebook
To use this model, you may install datasets and transformers:
Create a virtual environment:
python -m venv /path/to/venv
Activate the environment:
source /path/to/venv/bin/activate
Install the modules:
pip install datasets transformers
To transcribe audio in Catalan using this model, you can follow this example:
#Install Prerequisites
pip install torch
pip install datasets
pip install 'transformers[torch]'
pip install evaluate
pip install jiwer
#This code works with GPU
#Notice that: load_metric is no longer part of datasets.
# You have to remove it and use evaluate's load instead.
#(Note from November 2024)
import torch
from transformers import WhisperForConditionalGeneration, WhisperProcessor
#Load the processor and model.
MODEL_NAME="projecte-aina/whisper-large-v3-tiny-caesar"
processor = WhisperProcessor.from_pretrained(MODEL_NAME)
model = WhisperForConditionalGeneration.from_pretrained(MODEL_NAME).to("cuda")
#Load the dataset
from datasets import load_dataset, load_metric, Audio
ds=load_dataset("projecte-aina/3catparla_asr",split='test')
#Downsample to 16kHz
ds = ds.cast_column("audio", Audio(sampling_rate=16_000))
#Process the dataset
def map_to_pred(batch):
audio = batch["audio"]
input_features = processor(audio["array"], sampling_rate=audio["sampling_rate"], return_tensors="pt").input_features
batch["reference"] = processor.tokenizer._normalize(batch['normalized_text'])
with torch.no_grad():
predicted_ids = model.generate(input_features.to("cuda"))[0]
transcription = processor.decode(predicted_ids)
batch["prediction"] = processor.tokenizer._normalize(transcription)
return batch
#Do the evaluation
result = ds.map(map_to_pred)
#Compute the overall WER now.
from evaluate import load
wer = load("wer")
WER=100 * wer.compute(references=result["reference"], predictions=result["prediction"])
print(WER)
The specific dataset used to create the model is a corpus called CAESAR-tiny, which has not been released at the moment.
This model is the result of finetuning the model "openai/whisper-large-v3" by following this tutorial provided by Hugging Face.
If this model contributes to your research, please cite the work:
@misc{mena2024whisperlarge3catparla,
title={Acoustic Model in Catalan: whisper-large-v3-tiny-caesar.},
author={Hernandez Mena, Carlos Daniel; Giraldo, Jose ;Armentano-Oller, Carme; Solito, Sarah; Messaoudi, Abir; Costa, Federico; Zeballos, Rodolfo},
organization={Barcelona Supercomputing Center},
url={https://huggingface.co/projecte-aina/whisper-large-v3-tiny-caesar},
year={2024}
}
The fine-tuning process was performed during November (2024) in the Language Technologies Unit of the Barcelona Supercomputing Center by Carlos Daniel Hernández Mena.
For further information, please send an email to [email protected].
Copyright(c) 2024 by Language Technologies Unit, Barcelona Supercomputing Center.
This work has been promoted and financed by the Generalitat de Catalunya through the Aina project.
The training of the model was possible thanks to the computing time provided by Barcelona Supercomputing Center through MareNostrum 5.