Downloads · 30 days
7
2% of all-time downloads
CovianHive/next_bemba_ai_medium
next_bemba_ai_medium is a machine learning model from CovianHive. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<p align="center" <img src="https://huggingface.co/front/assets/huggingfacelogo-noborder.svg" height="80" / </p
Downloads · 30 days
7
2% of all-time downloads
All-time downloads
299
Public
Parameters
764M
15.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.1 GB · 100%
From the Hugging Face model README
Multilingual Whisper ASR (Automatic Speech Recognition) Fine-tuned Whisper model for Bemba and English using language tokens. Developed and maintained by NextInnoMind, led by Chalwe Silas.
WhisperForConditionalGeneration — fine-tuned using openai/whisper-medium
Framework: Transformers
Checkpoint Format: Safetensors
Languages: Bemba, English (with <|bem|> language token support)
This model is a Whisper Medium variant fine-tuned for Bemba and English, enabling robust multilingual transcription. It supports the use of language tokens (e.g., <|bem|>) to help guide decoding, particularly for low-resource languages like Bemba.
Base Model: openai/whisper-medium
Dataset:
Training Time: 8 epochs (~55 hours on A100 GPU)
Learning Rate: 1e-5
Batch Size: 16
Framework: Transformers + Accelerate
Tokenizer: WhisperProcessor with language="<|bem|>" and task="transcribe"
from transformers import pipeline
pipe = pipeline(
"automatic-speech-recognition",
model="NextInnoMind/next_bemba_ai_medium",
chunk_length_s=30,
return_timestamps=True
)
# Example
result = pipe("path_to_audio.wav")
print(result["text"])
📌 Tip: For Bemba, use the language token
<|bem|>to improve transcription accuracy.
<|bem|> is required for optimal Bemba decoding| Language | WER (Word Error Rate) | Dataset |
|---|---|---|
| Bemba | ~15.2% | BembaSpeech Eval Set |
| English | ~10.5% | Common Voice EN |
@misc{nextbembaai2025,
title={NextInnoMind next_bemba_ai_medium: Multilingual Whisper ASR model for Bemba and English},
author={Silas Chalwe and NextInnoMind},
year={2025},
howpublished={\url{https://huggingface.co/NextInnoMind/next_bemba_ai_medium}},
}
📬 Contact:
🔗 GitHub: SilasChalwe
Fine tuned in Zambia.