Downloads · 30 days
0
Sharanya186/TextToSpeech
TextToSpeech is a text-to-speech model from Sharanya186. Use it when you need text read aloud. It is set up for speechbrain. The card lists the license as apache-2.0.
This repository provides all the necessary tools for Text-to-Speech (TTS) with SpeechBrain using a Transformer pretrained on LJSpeech.
Downloads · 30 days
0
Access
Public
Updated Apr 21, 2024
Repo size
—
Likes
0
Public
Click a slice to open those files.
Other1.5 KB · 54%
From the Hugging Face model README
This repository provides all the necessary tools for Text-to-Speech (TTS) with SpeechBrain using a Transformer pretrained on LJSpeech.
The pre-trained model takes in text input and produces a spectrogram in output. One can get the final waveform by applying a vocoder (e.g., HiFIGAN) on top of the generated spectrogram.
import torchaudio
from speechbrain.inference.vocoders import HIFIGAN
texts = ["This is the example text"]
#initializing my model
my_tts_model = TextToSpeech.from_hparams(source="/content/")
#initializing vocoder(Hifigan) model
hifi_gan = HIFIGAN.from_hparams(source="speechbrain/tts-hifigan-ljspeech", savedir="tmpdir_vocoder")
# Running the TTS
mel_output = my_tts_model.encode_text(texts)
# Running Vocoder (spectrogram-to-waveform)
waveforms = hifi_gan.decode_batch(mel_output)
# Save the waverform
torchaudio.save('example_TTS.wav',waveforms.squeeze(1), 22050)