Downloads · 30 days
0
Cocii/glmtts-test
glmtts-test is a text-to-speech model from Cocii. Use it when you need text read aloud. The card lists the license as mit.
<div align="center" <a href="README.md" <img src="https://img.shields.io/badge/Language/语言-English-blue?style=flat-square" alt="English" </a <a href="READMEzh.md" <img src="https://img.shields.io/badge/Language/语言-中文-…
Downloads · 30 days
0
Access
Public
Updated Dec 10, 2025
Repo size
2.4 MB
Likes
0
Public
Click a slice to open those files.
.png2.4 MB · 99%
From the Hugging Face model README
<br><br>
<div align="center"> <img src="assets/images/logo.svg" width="50%"/> </div> <p align="center"> <a href="https://github.com/zai-org/GLM-TTS" target="_blank">💻 GitHub Repository</a> | <a href="https://huggingface.co/spaces/zai-org/GLM-TTS" target="_blank">🤗 Online Demo</a> | <a href="https://audio.z.ai/" target="_blank">🛠️ Audio.Z.AI</a> </p>GLM-TTS is a high-quality text-to-speech (TTS) synthesis system based on large language models, supporting zero-shot voice cloning and streaming inference. The system adopts a two-stage architecture combining an LLM for speech token generation and a Flow Matching model for waveform synthesis.
By introducing a Multi-Reward Reinforcement Learning framework, GLM-TTS significantly improves the expressiveness of generated speech, achieving more natural emotional control compared to traditional TTS systems.
GLM-TTS follows a two-stage design:
To tackle flat emotional expression, GLM-TTS uses a Group Relative Policy Optimization (GRPO) algorithm with multiple reward functions (Similarity, CER, Emotion, Laughter) to align the LLM's generation strategy.
Evaluated on seed-tts-eval. GLM-TTS_RL achieves the lowest Character Error Rate (CER) while maintaining high speaker similarity.
| Model | CER ↓ | SIM ↑ | Open-source |
|---|---|---|---|
| Seed-TTS | 1.12 | 79.6 | 🔒 No |
| CosyVoice2 | 1.38 | 75.7 | 👐 Yes |
| F5-TTS | 1.53 | 76.0 | 👐 Yes |
| GLM-TTS (Base) | 1.03 | 76.1 | 👐 Yes |
| GLM-TTS_RL (Ours) | 0.89 | 76.4 | 👐 Yes |
git clone [https://github.com/zai-org/GLM-TTS.git](https://github.com/zai-org/GLM-TTS.git)
cd GLM-TTS
pip install -r requirements.txt
python glmtts_inference.py \
--data=example_zh \
--exp_name=_test \
--use_cache \
# --phoneme # Add this flag to enable phoneme capabilities.
bash glmtts_inference.sh
We thank the following open-source projects for their support:
If you use GLM-TTS in your research, please cite:
@misc{glmtts2025,
title={GLM-TTS: Controllable & Emotion-Expressive Zero-shot TTS with Multi-Reward Reinforcement Learning},
author={CogAudio Group Members},
year={2025},
publisher={Zhipu AI Inc}
}