Downloads · 30 days
10
7% of all-time downloads
HiTZ/whisper-large-v3-ca
whisper-large-v3-ca is a machine learning model from HiTZ. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Whisper Large-V3 Catalan is an automatic speech recognition (ASR) model for Catalan speech. It is fine-tuned from [openai/whisper-large-v3] on the Catalan portion of Mozilla Common Voice 13.0, achieving a Word Error R…
Downloads · 30 days
10
7% of all-time downloads
All-time downloads
140
Public
Parameters
1.5B
6.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6.2 GB · 100%
From the Hugging Face model README
Whisper Large-V3 Catalan is an automatic speech recognition (ASR) model for Catalan speech. It is fine-tuned from [openai/whisper-large-v3] on the Catalan portion of Mozilla Common Voice 13.0, achieving a Word Error Rate (WER) of 5.97% on the Common Voice test split.
The model is intended for high-quality transcription of Catalan speech in a variety of accents and recording conditions, including read and semi-spontaneous speech.
This model leverages Whisper's multilingual pretraining and large-scale speech-text alignment, followed by supervised fine-tuning on Catalan speech data to improve language-specific accuracy.
Performance may degrade on:
The model may produce hallucinated text when audio quality is very poor or silent.
Biases present in the Common Voice dataset (e.g., demographic or accent imbalance) may be reflected in model outputs.
Users are encouraged to evaluate the model on their own data before deployment.
Dataset: Mozilla Common Voice 13.0 (Catalan subset)
Data type: Crowd-sourced, read speech
Preprocessing:
| Metric | Value |
|---|---|
| WER (test) | 5.97% |
These results indicate strong performance compared to the base Whisper multilingual model on Catalan speech.
| Training Loss | Epoch | Step | Validation Loss | Wer |
|---|---|---|---|---|
| 0.0988 | 1.95 | 1000 | 0.1487 | 6.5619 |
| 0.025 | 3.91 | 2000 | 0.1676 | 6.3155 |
| 0.0105 | 5.86 | 3000 | 0.1871 | 6.4035 |
| 0.0047 | 7.81 | 4000 | 0.1973 | 6.4870 |
| 0.0061 | 9.77 | 5000 | 0.2086 | 6.4836 |
| 0.0034 | 11.72 | 6000 | 0.2172 | 6.6442 |
| 0.0036 | 13.67 | 7000 | 0.2205 | 6.4041 |
| 0.002 | 15.62 | 8000 | 0.2214 | 6.4350 |
| 0.0011 | 17.58 | 9000 | 0.2339 | 6.1943 |
| 0.0009 | 19.53 | 10000 | 0.2388 | 6.2921 |
| 0.0011 | 21.48 | 11000 | 0.2327 | 6.2515 |
| 0.0003 | 23.44 | 12000 | 0.2472 | 6.2052 |
| 0.0012 | 25.39 | 13000 | 0.2382 | 6.2892 |
| 0.0001 | 27.34 | 14000 | 0.2550 | 5.9949 |
| 0.0006 | 29.3 | 15000 | 0.2574 | 6.3607 |
| 0.0001 | 31.25 | 16000 | 0.2584 | 6.0143 |
| 0.0001 | 33.2 | 17000 | 0.2686 | 5.9486 |
| 0.0 | 35.16 | 18000 | 0.2736 | 5.9194 |
| 0.0 | 37.11 | 19000 | 0.2768 | 5.9646 |
| 0.0 | 39.06 | 20000 | 0.2783 | 5.9714 |
from transformers import pipeline
hf_model = "HiTZ/whisper-large-v3-ca"
device = 0 # set to -1 for CPU
pipe = pipeline(
task="automatic-speech-recognition",
model=hf_model,
device=device
)
result = pipe("audio.wav")
print(result["text"])
If you use this model in your research, please cite:
@misc{dezuazo2025whisperlmimprovingasrmodels,
title={Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages},
author={Xabier de Zuazo and Eva Navas and Ibon Saratxaga and Inma Hernáez Rioja},
year={2025},
eprint={2503.23542},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
Please, check the related paper preprint in arXiv:2503.23542 for more details.
This model is available under the Apache-2.0 License. You are free to use, modify, and distribute this model as long as you credit the original creators.
For questions or issues, please open an issue in the model repository.