Downloads · 30 days
0
THU-SPMI/CTC-TTS
CTC-TTS is a text-to-speech model from THU-SPMI. Use it when you need text read aloud. It is set up for transformers. The card lists the license as apache-2.0.
[](https://arxiv.org/pdf/2602.19574) []() [](https://github.com/thu-spmi/CTC-TTS)
Downloads · 30 days
0
Access
Public
Updated Sep 15, 2026
Repo size
5.4 GB
Likes
0
Public
Click a slice to open those files.
.pt5.4 GB · 99%
From the Hugging Face model README
This repository contains official pre-trained checkpoints for CTC-TTS, proposed in our Interspeech 2026 paper CTC-TTS: LLM-Based Dual-Streaming Text-to-Speech with CTC Alignment.
CTC-TTS is a LLM-driven dual-streaming TTS framework, designed to solve two critical pain points of existing streaming synthesis systems:
We replace MFA with a CTC-based neural aligner (built on Whistle ASR model from THU-SPMI Lab) and propose a bi-word interleaving block structure. Two variants are provided for flexible quality-latency tradeoff:
Experiments on single-speaker streaming and multi-speaker zero-shot synthesis show consistent improvements over LLMVox, ELLA-V and other MFA-aligned baselines.
| Model Name | Params | Training Dataset |
|---|---|---|
| ctc-tts-l-singlespeaker | 34.7M | VoiceAssistant400K |
| ctc-tts-f-singlespeaker | 33.6M | VoiceAssistant400K |
| ctc-tts-l-multispeaker | 159.58M | LibriSpeech 960h |
| ctc-tts-f-multispeaker | 158.47M | LibriSpeech 960h |
Greedy search. FPL-A = first-packet latency assuming full text is available.
| Method | #Params | WER(↓) | CER(↓) | FPL-A ms(↓) | UTMOS(↑) |
|---|---|---|---|---|---|
| Ground Truth | NA | NA | NA | NA | 4.27 |
| LLMVox | 31.5M | 2.40 | 1.36 | 167 | 4.15 |
| CTC-TTS-F | 33.6M | 1.80 | 1.04 | 159 | 4.15 |
| CTC-TTS-L | 34.7M | 1.50 | 0.79 | 210 | 4.15 |
Given a text segment and its corresponding 3-second prefixed speech, synthesize speech for the remaining text. Nucleus sampling.
| Group | Method | #Params | WER(↓) | CER(↓) | SPK(↑) | UTMOS(↑) | MOS(↑) | SMOS(↑) |
|---|---|---|---|---|---|---|---|---|
| Ground Truth | — | NA | 1.92 | 0.69 | NA | 4.086 | 4.28±0.060 | 4.60±0.048 |
| Our Method | CTC-TTS-F | 158.47M | 5.20 | 2.68 | 0.930 | 4.013 | 4.31±0.057 | 4.58±0.050 |
| Our Method | CTC-TTS-L | 159.58M | 4.82 | 2.47 | 0.929 | 4.050 | 4.33±0.061 | 4.60±0.049 |
| Ablation | CTC+ELLA-V | 159.58M | 12.01 | 7.37 | 0.928 | 4.021 | 4.00±0.062 | 4.39±0.058 |
| Ablation | MFA+ELLA-V | 159.58M | 10.98 | 6.99 | 0.928 | 4.021 | 3.94±0.066 | 4.44±0.056 |
| Ablation | MFA+bi-word | 159.58M | 5.14 | 2.63 | 0.930 | 4.010 | 4.25±0.061 | 4.50±0.051 |
Given ~3 seconds of speech and its transcribed text as a prompt, synthesize speech for another utterance (out-of-domain). Nucleus sampling.
| Group | Method | #Params | WER(↓) | CER(↓) | SPK(↑) | UTMOS(↑) | MOS(↑) | SMOS(↑) |
|---|---|---|---|---|---|---|---|---|
| Ground Truth | — | NA | NA | NA | NA | 3.527 | 4.18±0.068 | 4.14±0.072 |
| Our Method | CTC-TTS-F | 158.47M | 8.02 | 4.20 | 0.880 | 3.903 | 4.16±0.064 | 3.85±0.071 |
| Our Method | CTC-TTS-L | 159.58M | 6.33 | 3.21 | 0.878 | 3.971 | 4.23±0.060 | 3.98±0.073 |
| Ablation | CTC+ELLA-V | 159.58M | 20.86 | 11.73 | 0.869 | 3.848 | 3.88±0.073 | 3.94±0.073 |
| Ablation | MFA+ELLA-V | 159.58M | 34.89 | 19.58 | 0.872 | 3.873 | 3.75±0.071 | 3.88±0.074 |
| Ablation | MFA+bi-word | 159.58M | 7.53 | 3.99 | 0.874 | 3.840 | 4.14±0.068 | 3.83±0.076 |
If you use this model or method in your research, please cite our paper:
@article{liu2026ctc,
title={CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment},
author={Liu, Hanwen and Yusuyin, Saierdaer and Huang, Hao and Ou, Zhijian},
journal={arXiv preprint arXiv:2602.19574},
year={2026}
}