Downloads · 30 days
19
11% of all-time downloads
Aleton/whisper-small-be-custom
whisper-small-be-custom is a automatic speech recognition model from Aleton. Use it when you need speech turned into text. The card lists the license as apache-2.0.
A fine-tuned version of openai/whisper-small optimized for Belarusian speech recognition. This model significantly outperforms both the base Whisper Small and even the much larger Whisper Large V3 on Belarusian speech.
Downloads · 30 days
19
11% of all-time downloads
All-time downloads
167
Public
Parameters
242M
4.8 GB on disk
Likes
2
Public
Click a slice to open those files.
.pt1.9 GB · 66%
From the Hugging Face model README
A fine-tuned version of openai/whisper-small optimized for Belarusian speech recognition. This model significantly outperforms both the base Whisper Small and even the much larger Whisper Large V3 on Belarusian speech.
| Model | Parameters | WER (lower is better) |
|---|---|---|
| openai/whisper-small (base) | 244M | 92.21% |
| openai/whisper-large-v3 | 1550M | 63.64% |
| This model | 244M | 20.21% |
Key finding: This fine-tuned Small model outperforms Whisper Large V3 by 36.37 percentage points, while being 6x smaller in size.
from transformers import pipeline
pipe = pipeline("automatic-speech-recognition", model="aleton/whisper-small-be-custom")
result = pipe("audio_file.mp3")
print(result["text"])
For longer audio files:
result = pipe("long_audio.mp3", chunk_length_s=30, batch_size=8)
print(result["text"])
This model demonstrates that targeted fine-tuning on language-specific data can dramatically improve performance for low-resource languages. The base Whisper models struggle with Belarusian due to limited representation in the original training data. Through fine-tuning on Common Voice 22 Sidon, this model achieves a 64.94 percentage point improvement over the base Small model and a 36.37 percentage point improvement over the Large V3 model.
Эта модель демонстрирует, что целенаправленное дообучение на языковых данных может значительно улучшить качество распознавания для малоресурсных языков. Базовые модели Whisper плохо справляются с белорусским языком из-за ограниченного представления в исходных обучающих данных. Благодаря дообучению на Common Voice 22 Sidon, эта модель показывает улучшение на 64.94 п.п. по сравнению с базовой Small моделью и на 36.37 п.п. по сравнению с Large V3.
Гэтая мадэль дэманструе, што мэтанакіраванае данавучанне на моўных дадзеных можа значна палепшыць якасць распазнавання для маларэсурсных моў. Базавыя мадэлі Whisper дрэнна спраўляюцца з беларускай мовай з-за абмежаванага прадстаўніцтва ў зыходных навучальных дадзеных. Дзякуючы данавучанню на Common Voice 22 Sidon, гэтая мадэль паказвае паляпшэнне на 64.94 п.п. у параўнанні з базавай Small мадэллю і на 36.37 п.п. у параўнанні з Large V3.
If you use this model, please cite:
@misc{whisper-small-be-custom,
author = {aleton},
title = {Whisper Small Belarusian Custom},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/aleton/whisper-small-be-custom}
}