Downloads · 30 days
300
66% of all-time downloads
SayedShaun/bengali-whisper-medium
bengali-whisper-medium is a automatic speech recognition model from SayedShaun. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
300
66% of all-time downloads
All-time downloads
452
Public
Parameters
764M
6.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.1 GB · 100%
From the Hugging Face model README
Bengali speech in, punctuated Bengali text out. A whisper-medium
fine-tune — the ASR half of the 1st-place solution of the
Bengali.AI Speech Recognition
Kaggle competition (0.312 WER, public leaderboard). Loads directly with
standard transformers tooling, no custom code required.
Attribution. Trained by tugstugi (Erdene-Ochir Tuguldur), team Chimege; redistributed here under Apache-2.0. This repo only adds packaging (safetensors conversion, a repaired generation config). Please cite the original author — see Citation.
Want punctuated output? Pair this with
asr-punctuation-restore-bn(same competition solution's MuRIL punctuation heads, published standalone — works on any ASR's output) via theasr-punct-restorepackage, used below.
Stage 1 — ASR:
from transformers import pipeline
asr = pipeline("automatic-speech-recognition",
model="SayedShaun/bengali-whisper-medium",
chunk_length_s=20.1) # >20s: competition audio is long-form
result = asr("clip.wav", generate_kwargs={"language": "bn", "task": "transcribe",
"num_beams": 4, "max_length": 260})
raw_text = result["text"] # no punctuation yet
Apply NFC normalization before comparing or storing output — the model emits precomposed Bengali letters (য় ড় ঢ়) that Bengali corpora write as base+nukta; they render identically but compare unequal (cost 18% WER on one otherwise-perfect sample):
import unicodedata
raw_text = unicodedata.normalize("NFC", raw_text)
Stage 2 — punctuation:
pip install git+https://github.com/sayedshaun/asr-punct-restore.git
from asr_punct_restore import PunctuationRestorer
restorer = PunctuationRestorer("SayedShaun/asr-punctuation-restore-bn",
layers=(12,)) # most accurate single head
punctuated = restorer(raw_text)
revision so a later push here can't change what you serve.।, ,, ? — no prosody, no exclamation points.Please cite the original author, not this packaging:
@misc{tuguldur2023bengaliasr,
author = {Tuguldur, Erdene-Ochir},
title = {1st place solution, Bengali.AI Speech Recognition},
year = {2023},
howpublished = {\url{https://www.kaggle.com/competitions/bengaliai-speech/writeups/chimege-1st-place-solution}}
}
tugstugi/bengali-ai-asr-submission on Kaggle, mirrored
at bengaliAI/tugstugi_bengaliai-asr_whisper-mediumApache-2.0, following the original upload; the Kaggle submission bundle is CC0.