Downloads · 30 days
14
7% of all-time downloads
omvishesh/SpeechT5_interview
SpeechT5_interview is a text-to-audio model from omvishesh. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
14
7% of all-time downloads
All-time downloads
204
Public
Parameters
144M
12.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors578 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of microsoft/speecht5_tts on an unknown dataset. It achieves the following results on the evaluation set:
The Speech T5 model is a text-to-speech (TTS) model based on the T5 architecture. It has been pretrained on a large corpus of speech data, allowing it to understand and generate human-like speech from input text. The model is capable of handling various speech synthesis tasks, making it suitable for applications such as virtual assistants, audiobook production, and more
More information needed
The model was trained using a custom-made dataset of 170 audio samples, containing commonly asked interview lines. Synthetic audio was generated using Amazon AWS Polly, which offered diverse voice options. The dataset was carefully curated to ensure a variety of speech styles, accents, and phonetic structures, enhancing the model's ability to generalize.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 0.6078 | 2.2535 | 40 | 0.4783 |
| 0.5393 | 4.5070 | 80 | 0.4533 |
| 0.4864 | 6.7606 | 120 | 0.4480 |
| 0.4846 | 9.0141 | 160 | 0.4493 |
| 0.4628 | 11.2676 | 200 | 0.4383 |
| 0.4731 | 13.5211 | 240 | 0.4392 |