Downloads · 30 days
0
0% of all-time downloads
SinhaSpeech/whisper-small-sinhala-experiments
whisper-small-sinhala-experiments is a automatic speech recognition model from SinhaSpeech. Use it when you need speech turned into text. The card lists the license as apache-2.0.
Experimental fine-tuned variants of openai/whisper-small for Sinhala automatic speech recognition (ASR). The best model, run11 (16.38% test WER), has its own repo: SinhaSpeech/whisper-small-sinhala-v6-e6-run11-best. A…
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
43
Public
Repo size
87.8 GB
Likes
0
Public
Click a slice to open those files.
.safetensors7.8 GB · 100%
From the Hugging Face model README
Experimental fine-tuned variants of openai/whisper-small for Sinhala automatic speech recognition (ASR). The best model, run11 (16.38% test WER), has its own repo: SinhaSpeech/whisper-small-sinhala-v6-e6-run11-best. All training data lives in SinhaSpeech/sinhala-asr-data.
This repo bundles ten training runs (run1 to run10) so the experiments are reproducible from one place.
models/
run1/ full fine-tune, v1 data — merged, ready-to-use model
run2/ LoRA (AMD hardware), v1 data — merged, ready-to-use model
run3/ LoRA adapter only, v1 data — best checkpoint (epoch 2), run interrupted at epoch 2.5
run4/ LoRA adapter only, v1 data — r=32, six target modules
run5/ full fine-tune, v4 data (speaker-disjoint, spacing-normalised)
run6/ full fine-tune, v5 data, lr 3e-5 linear, 5 epochs <- best v5 model
run7/ as run6 but cosine scheduler
run8/ as run6 but 4 epochs
run9/ as run6 but lr 2e-5 and 4 epochs
run10/ full fine-tune, v6 data, lr 3e-5 linear, 5 epochs (also SinhaSpeech/whisper-small-sinhala-v6-e5)
checkpoints/
run5/ resume notes and training log for run5
All splits (v1 to v6) are in SinhaSpeech/sinhala-asr-data under data/stratified_v1 to data/stratified_v6. The v1 to v4 copies that used to be bundled here were identical and have been removed.
Test WER / CER in percent. run1 to run4 were scored on the v1 test set (15,483 utterances), which is not speaker-disjoint, so its numbers are optimistic. run5 to run10 were scored on the speaker-disjoint test set (15,860 utterances); v5 also normalises about 8.5% of the labels, so run5 vs run6 is indicative only. Text is lower-cased with punctuation removed before scoring.
| Run | Type | Data | LR · scheduler | Epochs | Test WER | Test CER |
|---|---|---|---|---|---|---|
run1 | Full | v1 | 3e-5 · linear | 4 | 17.08 | 3.49 |
run2 | LoRA (merged) | v1 | 5e-5 · linear | 4 | 21.01 | 5.73 |
run3 | LoRA adapter | v1 | 3e-5 · cosine | interrupted at 2.5 | 59.07 | 18.03 |
run4 | LoRA adapter (r=32) | v1 | 1e-4 · cosine | 4 | 25.99 | 7.06 |
run5 | Full | v4 | 3e-5 · linear | 4 | 19.06 | 4.90 |
run6 | Full | v5 | 3e-5 · linear | 5 | 17.15 | 4.62 |
run7 | Full | v5 | 3e-5 · cosine | 5 | 17.44 | 4.72 |
run8 | Full | v5 | 3e-5 · linear | 4 | 17.90 | 4.76 |
run9 | Full | v5 | 2e-5 · linear | 4 | 19.07 | 5.02 |
run10 | Full | v6 | 3e-5 · linear | 5 | 17.17 | 4.71 |
| Run | Use when |
|---|---|
whisper-small-sinhala-v6-e6 (run11) | You want the best Sinhala accuracy (16.38% WER). Full weights, no PEFT dependency. |
run6 | The best v5-data model. |
run1 | You want the recipe that started the error analysis (v1 data). |
run2 | A merged LoRA model. |
run3, run4 | You want a small LoRA adapter to load on top of openai/whisper-small. |
run5, run7, run8, run9 | Comparisons: v4 data, cosine scheduler, fewer epochs, lower learning rate. |
run1, run2, run5 to run10)from transformers import WhisperForConditionalGeneration, WhisperProcessor
repo = "SinhaSpeech/whisper-small-sinhala-experiments"
model = WhisperForConditionalGeneration.from_pretrained(repo, subfolder="models/run6")
processor = WhisperProcessor.from_pretrained(repo, subfolder="models/run6")
run3, run4)from transformers import WhisperForConditionalGeneration, WhisperProcessor
from peft import PeftModel
from huggingface_hub import snapshot_download
base = WhisperForConditionalGeneration.from_pretrained("openai/whisper-small")
path = snapshot_download("SinhaSpeech/whisper-small-sinhala-experiments", allow_patterns="models/run3/*")
model = PeftModel.from_pretrained(base, f"{path}/models/run3")
processor = WhisperProcessor.from_pretrained(f"{path}/models/run3")
All runs were trained on Sinhala speech aggregated from OpenSLR-52, YouTube, BizBrains, and Linga sources (~154,828 examples total). See the dataset card for split details and column schema.
stratified (v1) and stratified_v2 may share speakers between train and eval splits; use stratified_v3 or later for a speaker-independent evaluation.no_repeat_ngram_size=3 is a recommended decoding setting for them.run5 to run10.