Downloads · 30 days
11
6% of all-time downloads
yhaha/EmoVoice
EmoVoice is a machine learning model from yhaha. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
<div align="center" <p align="center" <h1EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting</h1
Downloads · 30 days
11
6% of all-time downloads
All-time downloads
182
Public
Repo size
17.6 GB
Likes
3
Public
Click a slice to open those files.
.pt12.6 GB · 71%
From the Hugging Face model README
EmoVoice is a emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control. EmoVoice achieves SOTA performance on English EmoVoice-DB and Chinese Secap test sets.
<!-- ### Model <div align="center"> <img src="pics/emovoice_overview.png" alt="" width="500"> </div> ### Performance <table width="100%"> <tr> <td align="center"> <img src="pics/table2.png" alt="图片描述1" width="333"> </td> <td align="center"> <img src="pics/table3.png" alt="图片描述2" width="333"> </td> <td align="center"> <img src="pics/table4.png" alt="图片描述3" width="333"> </td> </tr> </table> --> <!-- ## Environmental Setup ```bash ### Create a separate environment if needed conda create -n EmoVoice python=3.10 conda activate EmoVoice pip install -r requirements.txt ``` ## Train and Inference ### Infer with checkpoints ```bash bash examples/tts/scripts/inference_EmoVoice.sh bash examples/tts/scripts/inference_EmoVoice-PP.sh bash examples/tts/scripts/inference_EmoVoice_1.5B.sh ``` ### Train from scratch ```bash # First Stage: Pretrain TTS bash examples/tts/scripts/pretrain_EmoVoice.sh bash examples/tts/scripts/pretrain_EmoVoice-PP.sh bash examples/tts/scripts/pretrain_EmoVoice_1.5B.sh # Second Stage: Finetune Emotional TTS bash examples/tts/scripts/ft_EmoVoice.sh bash examples/tts/scripts/ft_EmoVoice-PP.sh bash examples/tts/scripts/ft_EmoVoice_1.5B.sh ``` -->English model checkpoints of EmoVoice(0.5B), EmoVoice(1.5B) and EmoVoice-PP(0.5B) are uploaded.
Qwen2.5-0.5B-phn, the Qwen2.5-0.5B tokenizer with a phoneme-extended vocabulary, is uploaded.
If our work is useful for you, please cite as:
@article{yang2025emovoice,
title={EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting},
author={Yang, Guanrou and Yang, Chen and Chen, Qian and Ma, Ziyang and Chen, Wenxi and Wang, Wen and Wang, Tianrui and Yang, Yifan and Niu, Zhikang and Liu, Wenrui and others},
journal={arXiv preprint arXiv:2504.12867},
year={2025}
}
<!-- Paper link: https://arxiv.org/abs/2504.12867 -->
<!-- ## License
Our code is released under MIT License. The pre-trained models are licensed under the CC-BY-NC license due to the training data Emilia, which is an in-the-wild dataset. Sorry for any inconvenience this may cause.
-->