Downloads · 30 days
3
6% of all-time downloads
titaniumbones/whisper-base-mt-commonvoice
whisper-base-mt-commonvoice is a automatic speech recognition model from titaniumbones. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
openai/whisper-base fine-tuned for Maltese (mt) automatic speech recognition on Mozilla Common Voice Scripted Speech 26.0, trained on a laptop (Apple M3 Pro, PyTorch MPS backend, fp32).
Downloads · 30 days
3
6% of all-time downloads
All-time downloads
52
Public
Parameters
72.6M
290 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors290 MB · 98%
From the Hugging Face model README
openai/whisper-base fine-tuned for Maltese (mt) automatic speech recognition
on Mozilla Common Voice Scripted Speech 26.0, trained on a laptop
(Apple M3 Pro, PyTorch MPS backend, fp32).
| Model | Raw WER | Normalized WER | Test clips |
|---|---|---|---|
| whisper-base fine-tuned | 105.7% | 94.1% | 300 |
| whisper-base zero-shot | 107.6% | 134.7% | 300 |
Zero-shot Whisper barely functions on Maltese (≈1 h of it in the
original training mix); the WER delta above is what a few hours of
community-donated Common Voice speech and a laptop fine-tune buy. WER
computed on the official Common Voice test split (speaker-disjoint from
training data — note the training split draws on only 6
speakers, which limits generalization; see Limitations). "Normalized"
applies BasicTextNormalizer (lowercase, strip punctuation) to both
reference and hypothesis.
mozilla-foundation/common_voice_* datasets on this
Hub are empty shells).Standard Hugging Face Seq2SeqTrainer recipe
(fine-tune-whisper):
batch 4 × grad-accum 4, lr 1e-05, warmup 50,
500 steps, fp32 (MPS), gradient checkpointing, greedy decoding,
best-checkpoint-by-WER selection.
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="titaniumbones/whisper-base-mt-commonvoice")
print(asr("clip.mp3", generate_kwargs={"language": "maltese"}))