Downloads · 30 days
76
100% of all-time downloads
thantzinphyo/Whisper-Base-ASR
Whisper-Base-ASR is a automatic speech recognition model from thantzinphyo. Use it when you need speech turned into text. The card lists the license as apache-2.0.
This model is a fine-tuned version of openai/whisper-base for Automatic Speech Recognition (ASR) in Burmese .
Downloads · 30 days
76
100% of all-time downloads
All-time downloads
76
Public
Parameters
72.6M
290 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors290 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of openai/whisper-base for Automatic Speech Recognition (ASR) in Burmese .
my)The model was trained on a standardized, high-quality Burmese speech corpus:
\u1000-\u109F) with unified word segmentation and punctuation removal. The held-out test split evaluates zero-shot acoustic generalization across completely unseen voices.| Parameter | Value |
|---|---|
| Effective Batch Size | 64 (16 per device x 4 gradient accumulation) |
| Peak Learning Rate | 1.0e-4 |
| Learning Rate Scheduler | Cosine |
| Warmup Steps | 200 |
| Total Steps | 2,000 (~12.37 Epochs) |
| Mixed Precision | BF16 |
| Weight Decay | 0.01 |
| Gradient Checkpointing | Enabled |
| Augmentation | None |
| Stage / Evaluation | Step | Train Loss | Val Loss | WER (%) | CER (%) | SER (%) | DER (%) | IER (%) | chrF |
|---|---|---|---|---|---|---|---|---|---|
| Baseline (Untrained) | 0 | - | - | 753.85 | 437.94 | 100.00 | 7.69 | 653.85 | 0.00 |
| Validation Set | 250 | 0.6374 | 0.0725 | 47.99 | 10.33 | 93.57 | 10.36 | 2.65 | 81.45 |
| Validation Set | 500 | 0.2351 | 0.0453 | 37.15 | 6.92 | 87.52 | 6.12 | 2.81 | 86.66 |
| Validation Set | 750 | 0.1369 | 0.0417 | 31.71 | 5.70 | 81.61 | 3.60 | 4.19 | 89.08 |
| Validation Set | 1000 | 0.0482 | 0.0416 | 28.87 | 5.00 | 76.74 | 4.22 | 3.29 | 91.07 |
| Validation Set | 1250 | 0.0252 | 0.0435 | 28.18 | 4.79 | 76.74 | 3.72 | 3.31 | 91.23 |
| Validation Set | 1500 | 0.0039 | 0.0476 | 28.06 | 4.69 | 78.48 | 3.81 | 2.94 | 91.40 |
| Validation Set | 1750 | 0.0007 | 0.0502 | 26.91 | 4.47 | 76.30 | 3.26 | 3.44 | 91.79 |
| Validation Set | 2000 | 0.0004 | 0.0507 | 26.55 | 4.43 | 75.65 | 3.31 | 3.37 | 91.94 |
| UNSEEN TEST (Final) | Final | - | - | 36.29 | 8.55 | 94.30 | 5.40 | 3.22 | 83.71 |
import torch
from transformers import pipeline
pipe = pipeline(
'automatic-speech-recognition',
model='thantzinphyo/Whisper-Base-ASR',
device='cuda:0' if torch.cuda.is_available() else 'cpu',
)
output = pipe('audio.wav', generate_kwargs={'language': 'burmese', 'task': 'transcribe'})
print(output['text'])
Apache-2.0