Downloads · 30 days
522
92% of all-time downloads
MCAA1-MSU/TTSModel
TTSModel is a text-to-speech model from MCAA1-MSU. Use it when you need text read aloud. The card lists the license as cc-by-nc-4.0.
A Swahili text-to-speech model, finetuned from Meta's MMS-TTS Swahili checkpoint on the swatts split of google/WaxalNLP, using the VITS finetuning recipe from ylacombe/finetune-hf-vits. Developed by the Maseno Centre…
Downloads · 30 days
522
92% of all-time downloads
All-time downloads
569
Public
Parameters
36.3M
145 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors145 MB · 100%
From the Hugging Face model README
A Swahili text-to-speech model, finetuned from Meta's MMS-TTS Swahili checkpoint on the swa_tts split of google/WaxalNLP, using the VITS finetuning recipe from ylacombe/finetune-hf-vits. Developed by the Maseno Centre for Applied AI (MCAAI).
facebook/mms-tts-swh (Meta's Massively Multilingual Speech TTS, Swahili)swa_tts config — 1,387 train utterances (805 after duration/length filtering), single Swahili speaker, 16kHz audio, sourced via the Loud & Clear initiative.swh / ISO 639-3)| Setting | Value |
|---|---|
| Learning rate | 2e-5 |
| Batch size | 8 |
| Precision | fp16 |
| Max clip duration | 20s |
| Min clip duration | 0.5s |
| Loss weights | mel=35, kl=1.5, disc=3, gen/fmaps/duration=1 |
Training used the finetune-hf-vits recipe with transformers==4.35.1, datasets==2.14.7, accelerate==0.24.1, numpy<2.0.
import numpy as np
from transformers import pipeline
from IPython.display import Audio as IPyAudio
synthesiser = pipeline("text-to-speech", model="MCAA1-MSU/TTSModel")
speech = synthesiser("Habari yako, karibu Kenya.")
audio = np.squeeze(speech["audio"])
display(IPyAudio(audio, rate=speech["sampling_rate"]))
Verified loading cleanly with transformers pipeline("text-to-speech", ...) — no missing/unexpected weight warnings on load.
Research and experimentation with Swahili TTS.
Derived from facebook/mms-tts-swh (CC-BY-NC-4.0, non-commercial). This finetuned model inherits that license. swa_tts training data is CC-BY-SA-4.0.