Downloads · 30 days
63
31% of all-time downloads
DigitalUmuganda/lingala_vits_tts
lingala_vits_tts is a machine learning model from DigitalUmuganda. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This model was trained on the OpenSLR's 71.6 hours aligned lingala bible dataset.
Downloads · 30 days
63
31% of all-time downloads
All-time downloads
206
Public
Repo size
373 MB
Likes
2
Public
Click a slice to open those files.
.pth373 MB · 100%
From the Hugging Face model README
This model was trained on the OpenSLR's 71.6 hours aligned lingala bible dataset.
A Conditional Variational Autoencoder with Adversarial Learning(VITS), which is an end-to-end approach to the text-to-speech task. To train the model, we used the espnet2 toolkit.
First install espnet2
pip install espnet
Download the model and the config files from this repo. To generate a wav file using this model, run the following:
from espnet2.bin.tts_inference import Text2Speech
import soundfile as sf
text2speech = Text2Speech(train_config="config.yaml",model_file="train.total_count.best.pth")
wav = text2speech("oyo kati na Ye ozwi lisiko mpe bolimbisi ya masumu")["wav"]
sf.write("outfile.wav", wav.numpy(), text2speech.fs, "PCM_16")