Downloads · 30 days
0
nmcuong/MeloTTS-Vietnamese
MeloTTS-Vietnamese is a text-to-speech model from nmcuong. Use it when you need text read aloud. The card lists the license as mit.
<div align="center" <div </div <img src="logo.png" width="300"/ <br <a href="https://trendshift.io/repositories/8133" target="blank"<img src="https://trendshift.io/api/badge/repositories/8133" alt="myshell-ai%2FM…
Downloads · 30 days
0
Access
Public
Updated Feb 26, 2026
Repo size
1.2 GB
Likes
7
Public
Click a slice to open those files.
.pth1.2 GB · 100%
From the Hugging Face model README
MeloTTS is a high-quality, open-source text-to-speech system developed by MyShell AI. It is built on top of the VITS/VITS2 architecture and uses BERT-based linguistic features to produce natural-sounding speech. MeloTTS supports multiple languages and is designed to be fast enough for real-time CPU inference.
Strengths of the original MeloTTS:
Limitations of the original MeloTTS:
MeloTTS Vietnamese is a version of MeloTTS specifically optimized for the Vietnamese language. It inherits the high-quality and fast-inference characteristics of the original model while introducing targeted improvements to handle the unique phonological properties of Vietnamese — including its 6 tones, complex vowel system, and syllable structure.
This model is designed to produce natural, accurate Vietnamese speech and can be easily fine-tuned on custom Vietnamese datasets.
melo/text/symbols.pyThis model was fine-tuned from the base MeloTTS model by:
The pre-trained model can be downloaded from Hugging Face:
git clone https://github.com/manhcuong02/MeloTTS_Vietnamese.git
cd MeloTTS_Vietnamese
pip install -r requirements.txt
Download the model checkpoint and config from Hugging Face and place them in your desired directory.
Refer to the notebook test_infer.ipynb for a full example. Basic usage:
from melo.api import TTS
# Speed is adjustable
speed = 1.0
# You can set device to 'cpu', 'cuda', 'cuda:0', or 'mps'
device = "cuda:0" # Will automatically use GPU if available
# Load the Vietnamese TTS model
model = TTS(
language="VI",
device=device,
config_path="/path/to/config.json",
ckpt_path="/path/to/G_model.pth",
)
speaker_ids = model.hps.data.spk2id
# Convert text to speech
text = "Nhập văn bản tại đây"
output_path = "output.wav"
model.tts_to_file(text, speaker_ids["speaker_name"], output_path, speed=speed, quiet=True)
The full data preparation process is detailed in docs/training.md. At minimum, you need:
path/to/audio_001.wav |<speaker_name>|<language_code>|<text_001>
path/to/audio_002.wav |<speaker_name>|<language_code>|<text_002>
Run the preprocessing script to prepare training data:
python melo/preprocess_text.py \
--metadata /path/to/text_training.list \
--config_path /path/to/config.json \
--device cuda:0 \
--val-per-spk 10 \
--max-val-total 500
Alternatively, use the shell script melo/preprocess_text.sh with appropriate parameters.
Follow the training instructions in docs/training.md.
The Vietnamese adaptation, code implementation, and fine-tuning of this model were developed by Nguyễn Mạnh Cường.
Listen to sample outputs from the model:
"Buổi sáng ở thành phố bắt đầu bằng tiếng xe cộ nhộn nhịp và ánh nắng nhẹ xuyên qua những tòa nhà cao tầng."
<audio controls src="https://huggingface.co/nmcuong/MeloTTS_Vietnamese/resolve/main/samples/sample.wav"></audio>
"Người đi làm vội vã, học sinh ríu rít trò chuyện, còn quán cà phê góc phố thì thoang thoảng mùi thơm dễ chịu."
<audio controls src="https://huggingface.co/nmcuong/MeloTTS_Vietnamese/resolve/main/samples/sample-2.wav"></audio>
"Cuối cùng, hãy thử thì thầm một câu thật nhẹ nhàng, rồi bất ngờ chuyển sang giọng nói to, rõ và đầy năng lượng."
<audio controls src="https://huggingface.co/nmcuong/MeloTTS_Vietnamese/resolve/main/samples/sample-3.wav"></audio>
This project is licensed under the MIT License, consistent with the original MeloTTS project. It may be used for both commercial and non-commercial purposes.
This implementation is based on TTS, VITS, VITS2, and Bert-VITS2. We appreciate their outstanding work.