Downloads · 30 days
27
17% of all-time downloads
thunderboltc/whisper-small-santali-coded
whisper-small-santali-coded is a automatic speech recognition model from thunderboltc. Use it when you need speech turned into text. The card lists the license as apache-2.0.
A full fine-tune of openai/whisper-small for automatic speech recognition on Santali, transcribed into Sanlish — a Latin-script romanization scheme for Santali speech.
Downloads · 30 days
27
17% of all-time downloads
All-time downloads
160
Public
Parameters
242M
16.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors967 MB · 100%
From the Hugging Face model README
A full fine-tune of openai/whisper-small for automatic speech recognition on Santali, transcribed into Sanlish — a Latin-script romanization scheme for Santali speech.
Unlike partial-freeze / adapter-style approaches, every parameter of the base Whisper-small checkpoint (encoder, decoder, and output head) was fine-tuned. No language tag was pinned during tokenization or generation (generation_config.language = None), so the model was not steered toward Bengali or any other Whisper-supported language — it learned Sanlish surface forms purely from the fine-tuning data.
openai/whisper-small (~244M params)santali_train.csv, santali_val.csv, santali_test.csv)audio_path (16kHz audio) → sanlish (target transliteration text)| Hyperparameter | Value |
|---|---|
| Base checkpoint | openai/whisper-small |
| Learning rate | 2e-5 |
| Weight decay | 0.1 |
| Batch size (train) | 16 |
| Epochs (planned / completed) | 20 planned, 17 completed (training interrupted by a runtime disconnect) |
| Warmup steps | computed dynamically as 10% of total training steps |
| Eval/save strategy | every epoch |
| Early stopping | patience 5 (on validation WER) — not yet triggered when training stopped |
| Metric for best model | WER (lower is better) |
Per-epoch validation metrics (normalized WER/CER — lowercased, punctuation-stripped):
| Epoch | Train Loss | Val Loss | WER (%) | CER (%) |
|---|---|---|---|---|
| 1 | 1.8145 | 0.6641 | 54.51 | 11.61 |
| 3 | 0.1810 | 0.4263 | 36.37 | 6.60 |
| 8 | 0.0142 | 0.4845 | 33.30 | 5.92 |
| 10 | 0.0068 | 0.4903 | 31.50 | 5.58 |
| 14 (best WER) | 0.0006 | 0.5197 | 31.07 | 5.39 |
17 (checkpoint on main) | 0.0003 | 0.5266 | 31.18 | 5.41 |
Note: validation loss bottoms out around epoch 3 and rises afterward while WER/CER keep improving slightly — a sign of overfitting on the loss objective that hasn't yet hurt transcription accuracy. Epoch 14 had the lowest validation WER of the run; epoch 17 (the last checkpoint pushed before the training runtime disconnected) is marginally behind it and is what's currently on main. Earlier epoch checkpoints are available in this repo's commit history.
Test set: 28.82% WER, 4.89% CER — evaluated on the held-out test split (193 examples), using the epoch-17 checkpoint currently on main.
task="transcribe" at inference and not rely on Whisper's automatic language detection for this checkpoint.from transformers import WhisperForConditionalGeneration, WhisperProcessor
model = WhisperForConditionalGeneration.from_pretrained("thunderboltc/whisper-small-santali-coded")
processor = WhisperProcessor.from_pretrained("thunderboltc/whisper-small-santali-coded", task="transcribe")
model.generation_config.language = None
model.generation_config.task = "transcribe"
# input_features = processor(audio_array, sampling_rate=16000, return_tensors="pt").input_features
# predicted_ids = model.generate(input_features)
# transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)