Downloads Β· 30 days
20
24% of all-time downloads
Ai4Innov/SALAMA_SM_ASR
SALAMA_SM_ASR is a automatic speech recognition model from Ai4Innov. Use it when you need speech turned into text. The card lists the license as apache-2.0.
Developer: AI4NNOV Authors: AI4NNOV. Version: v1.0 License: Apache 2.0 Model Type: Automatic Speech Recognition (ASR) Base Model: openai/whisper-small (fine-tuned for Swahili)
Downloads Β· 30 days
20
24% of all-time downloads
All-time downloads
83
Public
Parameters
242M
11.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt7.7 GB Β· 61%
From the Hugging Face model README
Developer: AI4NNOV
Authors: AI4NNOV.
Version: v1.0
License: Apache 2.0
Model Type: Automatic Speech Recognition (ASR)
Base Model: openai/whisper-small (fine-tuned for Swahili)
SALAMA-STT (Speech-to-Text) is the first module of the SALAMA Framework β a modular end-to-end speech-to-speech AI system built for African languages.
This model is fine-tuned from OpenAIβs Whisper-small architecture for Swahili speech recognition, enhancing performance on African accents and conversational data.
The model converts Swahili audio input into accurate transcriptions and serves as the entry point for downstream LLM and TTS modules.
SALAMA-STT leverages the Whisper-small architecture with a transformer encoder-decoder optimized for low-resource Swahili audio transcription tasks.
The model was fine-tuned on the Mozilla Common Voice 17.0 Swahili dataset, ensuring robustness to diverse accents and speech clarity.
| Parameter | Value |
|---|---|
| Base Model | openai/whisper-small |
| Fine-Tuning | Full model fine-tuning (fp16 precision) |
| Optimizer | AdamW |
| Learning Rate | 1e-5 |
| Batch Size | 16 |
| Epochs | 10 |
| Frameworks | Transformers + Datasets + TorchAudio |
| Languages | Swahili (sw), English (en) |
| Dataset | Description | Purpose |
|---|---|---|
mozilla-foundation/common_voice_17_0 | 20 hours of Swahili speech data | Supervised fine-tuning |
| Custom local Swahili recordings | Conversational + accent-rich data | Accent robustness |
| Common Voice validation split | 2.3 hours | Evaluation |
| Metric | Baseline (Whisper-small) | Fine-tuned (SALAMA-STT) | Improvement |
|---|---|---|---|
| WER (Word Error Rate) | 1.15 | 0.43 | π» 62% |
| CER (Character Error Rate) | 0.39 | 0.18 | π» 54% |
| Accuracy | 85.2% | 95.4% | +10.2% |
Evaluation conducted on a 2-hour held-out Swahili validation set from Common Voice.
Below is a quick example for Swahili speech transcription using this model:
from transformers import pipeline
# Load Swahili Whisper ASR
asr_pipeline = pipeline(
"automatic-speech-recognition",
model="EYEDOL/salama-stt",
chunk_length_s=30,
device_map="auto"
)
# Example audio file (replace with your file)
audio_path = "swahili_audio_sample.wav"
# Transcribe audio
result = asr_pipeline(audio_path)
print("π£οΈ Transcription:")
print(result["text"])
Example Output:
βKaribu kwenye mfumo wa SALAMA unaosaidia kutambua na kuelewa sauti ya Kiswahili kwa usahihi mkubwa.β
| Dataset | Metric | Score |
|---|---|---|
| Common Voice 17.0 (test) | WER | 0.43 |
| Common Voice 17.0 (test) | CER | 0.18 |
| Local Swahili Test Set | Accuracy | 95.4% |
| Model | Description |
|---|---|
EYEDOL/salama-llm | Swahili instruction-tuned LLM for reasoning and dialogue |
EYEDOL/salama-tts | Swahili text-to-speech (VITS) model for natural speech synthesis |