Downloads · 30 days
10
12% of all-time downloads
Professor/whisper-medium-afrispeech
whisper-medium-afrispeech is a automatic speech recognition model from Professor. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as cc-by-nc-sa-4.0.
Fine-tune of openai/whisper-medium on AfriSpeech-200 for African-accented English speech recognition, spanning both general and clinical/medical domains.
Downloads · 30 days
10
12% of all-time downloads
All-time downloads
84
Public
Parameters
764M
3.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.1 GB · 100%
From the Hugging Face model README
Fine-tune of openai/whisper-medium on
AfriSpeech-200 for African-accented English speech recognition, spanning both
general and clinical/medical domains.
Zero-shot = openai/whisper-medium evaluated on the same test set with no fine-tuning.
| Metric | Zero-shot (baseline) | Fine-tuned (this model) | Improvement |
|---|---|---|---|
| Overall WER | 43.22% | 20.13% | −23.1 pts (−53%) |
| Clinical WER | 50.55% | 27.47% | −23.1 pts (−46%) |
| General WER | 36.09% | 12.98% | −23.1 pts (−64%) |
Fine-tuning more than halves WER across the board.
| Domain | WER | n |
|---|---|---|
| General | 12.98% | 2,670 |
| Clinical | 27.47% | 3,508 |
| Overall | 20.13% | 6,178 |
Clinical speech remains ~2× harder than general — medical terminology (drug names, conditions, dosages) drives most of the error, so domain matters when reporting WER.
Performance varies widely across accents (~7% to ~45%):
| Best accents | WER | Hardest accents | WER | |
|---|---|---|---|---|
| okirika | 7.4% | khana | 44.6% | |
| afrikaans | 8.0% | nyandang | 44.1% | |
| brass | 8.6% | |||
| ikwere | 8.9% | |||
| twi | 10.3% |
The hardest accents are mostly smaller, under-represented ones — an important
coverage/equity consideration. Full per-accent numbers: eval_wer_breakdown.csv.
import torch
from transformers import WhisperForConditionalGeneration, WhisperProcessor
model = WhisperForConditionalGeneration.from_pretrained("Professor/whisper-medium-afrispeech")
processor = WhisperProcessor.from_pretrained("Professor/whisper-medium-afrispeech")
# audio: a 16 kHz mono waveform (numpy array)
inputs = processor(audio, sampling_rate=16000, return_tensors="pt")
ids = model.generate(inputs.input_features, language="english", task="transcribe")
print(processor.batch_decode(ids, skip_special_tokens=True)[0])
Trained on AfriSpeech-200 by Intron Health, released under CC-BY-NC-SA-4.0 (non-commercial, share-alike, attribution). This derivative carries the same license.
@article{olatunji2023afrispeech,
title={AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR},
author={Olatunji, Tobi and others},
journal={Transactions of the Association for Computational Linguistics},
year={2023}
}