Downloads ยท 30 days
22
30% of all-time downloads
Bird/ThonburianTTS
ThonburianTTS is a text-to-speech model from Bird. Use it when you need text read aloud. It is set up for f5-tts. The card lists the license as cc.
<p align="center" <img src="assets/ThonburianTTSLogo.png" width="400"/<br <img src="assets/looloo-logo.png" width="150" / </p
Downloads ยท 30 days
22
30% of all-time downloads
All-time downloads
73
Public
Repo size
4 GB
Likes
0
Public
Click a slice to open those files.
.safetensors4 GB ยท 100%
From the Hugging Face model README
๐ Model Checkpoints | ๐ค Gradio Demo | ๐ ThonburianTTS Paper | Colab Notebook | GitHub
Thonburian TTS is a Thai Text-to-Speech (TTS) engine built on top of the F5-TTS.
It generates natural and expressive Thai speech by leveraging Flow-Matching diffusion techniques and can mimic reference voices from short audio samples. The system supports:
language="th")| Model Component | Description | URL |
|---|---|---|
| F5-TTS Thai | Flow Matching-based Thai TTS models | Link |
| F5-TTS IPA | Flow Matching-based Thai-IPA TTS models | Link |
Install dependencies:
pip install torch cached-path librosa transformers f5-tts
sudo apt install ffmpeg
git clone https://github.com/biodatlab/thonburian-tts.git
cd thonburian-tts
from flowtts.inference import FlowTTSPipeline, ModelConfig, AudioConfig
import torch
# Configure F5-TTS model
model_config = ModelConfig(
language="th",
model_type="F5",
checkpoint="hf://biodatlab/ThonburianTTS/megaF5/mega_f5_last.safetensors",
vocab_file="hf://biodatlab/ThonburianTTS/megaF5/mega_vocab.txt",
vocoder="vocos",
device="cuda" if torch.cuda.is_available() else "cpu"
)
# Basic audio settings
audio_config = AudioConfig(
silence_threshold=-45,
cfg_strength=2.5,
speed=1.0
)
pipeline = FlowTTSPipeline(model_config, audio_config)
from flowtts.inference import FlowTTSPipeline, ModelConfig, AudioConfig
import torch
# Configure F5-TTS model
model_config = ModelConfig(
model_type="F5",
checkpoint="hf://biodatlab/ThonburianTTS/megaIPA/model_last_prune.safetensors",
vocab_file="hf://biodatlab/ThonburianTTS/megaIPA/mega_vocab_ipa.txt",
vocoder="vocos",
device="cuda" if torch.cuda.is_available() else "cpu"
)
# Basic audio settings
audio_config = AudioConfig(
silence_threshold=-45,
cfg_strength=2.5,
speed=1.0
)
pipeline = FlowTTSPipeline(model_config, audio_config)
If you use ThonburianTTS in your research, please cite:
@INPROCEEDINGS{11320472,
author={Aung, Thura and Sriwirote, Panyut and Thavornmongkol, Thanachot and Pipatsrisawat, Knot and Achakulvisut, Titipat and Aung, Zaw Htet},
booktitle={2025 20th International Joint Symposium on Artificial Intelligence and Natural Language Processing (iSAI-NLP)},
title={ThonburianTTS: Enhancing Neural Flow Matching Models for Authentic Thai Text-to-Speech},
year={2025},
volume={},
number={},
pages={1-6},
keywords={Adaptation models;Codes;Accuracy;Error analysis;Phonetics;Robustness;Natural language processing;Text to speech;Noise measurement;Research and development;Thai text-to-speech;Flow matching;F5-TTS},
doi={10.1109/iSAI-NLP66160.2025.11320472}}
Thura Aung, Panyut Sriwirote, Thanachot Thavornmongkol, Knot Pipatsrisawat, Titipat Achakulvisut, Zaw Htet Aung, "ThonburianTTS: Enhancing Neural Flow Matching Models for Authentic Thai Text-to-Speech", 2025 20th International Joint Symposium on Artificial Intelligence and Natural Language Processing (iSAI-NLP), Phuket, Thailand, 2025, pp. 1-6, doi: 10.1109/iSAI-NLP66160.2025.11320472.
The models are released under the Creative Commons Attribution Non-Commercial ShareAlike 4.0 License (CC BY-NC-SA 4.0).
We would like to acknowledge NSTDA Supercomputer Center (ThaiSC) project #pv824003 for providing computing resources for this work.