Downloads · 30 days
34
42% of all-time downloads
RobotsMali/bam-vits
bam-vits is a text-to-audio model from RobotsMali. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-4.0.
RobotsMali/bam-vits is an experimental multi-speaker Bambara (Bamanankan, bm) text-to-speech checkpoint from RobotsMali AI4D Lab. It adapts VITS from ylacombe/vits-vctk-with-discriminator, an English VCTK checkpoint,…
Downloads · 30 days
34
42% of all-time downloads
All-time downloads
81
Public
Parameters
39.6M
159 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors159 MB · 100%
From the Hugging Face model README
RobotsMali/bam-vits is an experimental multi-speaker Bambara (Bamanankan, bm) text-to-speech checkpoint from RobotsMali AI4D Lab. It adapts VITS from ylacombe/vits-vctk-with-discriminator, an English VCTK checkpoint, to Bambara.
Research checkpoint — substantially undertrained. This was produced as RobotsMali's first TTS experiment. It was trained for 200 epochs on a small, noisy subset. VITS systems are normally trained for hundreds of thousands of optimizer steps; RobotsMali uses roughly 200,000 steps as a useful reference budget. This run is far below that scale. Expect unstable pronunciation, noise, unnatural prosody, speaker leakage, and occasional unintelligible output. No objective or human evaluation is available.
transformers.VitsModel; the discriminator was removedThe project upgrades the upstream dependencies while keeping its core training approach largely unchanged.
pip install "transformers>=5.9" torch soundfile
import soundfile as sf
import torch
from transformers import AutoTokenizer, VitsModel
repo_id = "RobotsMali/bam-vits"
text = "An ka taa sugu la."
speaker_id = 0 # valid range: 0 through 19
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = VitsModel.from_pretrained(repo_id)
inputs = tokenizer(text.lower(), return_tensors="pt")
with torch.no_grad():
waveform = model(**inputs, speaker_id=speaker_id).waveform[0]
sf.write("bambara.wav", waveform.cpu().numpy(), model.config.sampling_rate)
The training script remapped participant identifiers to contiguous indices without exporting that mapping. Model IDs 0–19 therefore cannot be reliably mapped back to named AfVoices participants. Audition them for research; do not present an ID as a verified identity.
The model used top20-speakers from RobotsMali/afvoices-notag: 21,253 train and 1,128 test examples from the 20 AfVoices participants with the most utterances, excluding transcripts with semantic/acoustic tags. AfVoices is spontaneous, variably noisy speech collected for ASR rather than studio TTS.
The reproducible configuration is config/bam-vits.yaml.
| Setting | Value |
|---|---|
| Epochs | 200 |
| Per-device train batch size | 80 |
| Learning rate | 0.0005 |
| Precision | FP16 |
| Audio duration filter | 0.2–20 s |
| Maximum token length | 450 |
| Seed | 789 |
Vocabulary and speaker embeddings were resized. Text was lowercased and audio resampled to 22.05 kHz.
No MOS, intelligibility, pronunciation, speaker-similarity, or automated TTS results are reported. The held-out split was used during training, but losses are not perceptual evaluation.
Compared informally with bam-vits-pseudo-ipa, pseudo-IPA produced no remarkable improvement in quality or convergence; plain Bambara often sounded slightly more natural. RobotsMali's working explanation is that Bambara orthography is already largely phonetic and that the large English–Bambara acoustic mismatch gave the transferred waveform generator little useful English-phonetic guidance. This is an observation from an undertrained experiment, not a controlled conclusion.
This release is for low-resource TTS research, listening experiments, baselines, and further fine-tuning. It is not production quality. The limited, non-studio corpus can cause noise, disfluencies, pronunciation errors, demographic imbalance, and speaker leakage. Numbers, abbreviations, foreign words, code-switching, unusual punctuation, long text, and non-Bambara input may fail. Safety, bias, memorization, and voice similarity have not been systematically evaluated.
Do not use it where intelligibility or identity is safety-critical, to impersonate a person, or to create deceptive audio. Use short Bambara sentences, listen critically, and disclose that output is synthetic and experimental.
bam-vits-train: discriminator retained for continued trainingbam-vits-fintech: adapted on single-speaker FinBamSpeechbam-vits-pseudo-ipa: pseudo-IPA experimentPlease cite VITS and identify this checkpoint by repository ID.
@inproceedings{kim2021vits,
title={Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech},
author={Kim, Jaehyeon and Kong, Jungil and Son, Juhee},
booktitle={Proceedings of the 38th International Conference on Machine Learning},
pages={5530--5540},
year={2021}
}
Questions are welcome in the project repository or this model's Community tab.