Downloads · 30 days
0
botp/WordVoice-base-0.5B
WordVoice-base-0.5B is a text-to-speech model from botp. Use it when you need text read aloud. The card lists the license as apache-2.0.
WordVoice is a novel speech generation framework that transforms traditional implicit end-to-end TTS generation into an explicit, highly controllable paradigm. Built on top of the CosyVoice3 framework, it enables prec…
Downloads · 30 days
0
Access
Public
Updated Jul 24, 2026
Repo size
5.4 GB
Likes
0
Public
Click a slice to open those files.
.pt5.4 GB · 100%
From the Hugging Face model README
WordVoice is a novel speech generation framework that transforms traditional implicit end-to-end TTS generation into an explicit, highly controllable paradigm. Built on top of the CosyVoice3 framework, it enables precise and decoupled word-level control over five acoustic dimensions.
Supports independent and decoupled control of five acoustic attributes for each input word:
b0–b4).0–1).-1–1).Employs a bound-token (<b>) mechanism within the autoregressive (AR) language model. Before generating the speech tokens for a specific word, the model explicitly predicts its acoustic attributes, realizing an intelligent process of "planning prosody first, then generating sound."
We recommend using Conda to manage your Python environment.
conda create -n wordvoice python=3.10 -y
conda activate wordvoice
git clone https://github.com/XXH333/WordVoice-main.git
cd WordVoice-main
pip install -e .
pip install num2words==0.5.14 x_transformers==2.11.24
Run the following script to automatically download the pre-trained weights and dependencies (such as CosyVoice3, MMS-FA, etc.):
bash download_models.sh
You can run the out-of-the-box inference script to experience both the Free Mode and Control Mode of WordVoice:
python wordvoice_infer.py
For custom prompts and detailed control parameters, refer to wordvoice_infer.py and the GitHub repository.
If you find this work or the models useful, please cite:
@misc{nie2026wordvoice,
title={WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS},
author={Sihang Nie and Jinxin Ji and Xiaofen Xing and Deyi Tuo and Chengbin Jin and Jialong Mai and Xiangmin Xu},
year={2026},
eprint={2607.06461},
archivePrefix={arXiv},
primaryClass={eess.AS},
url={https://arxiv.org/abs/2607.06461},
}