Downloads · 30 days
9
100% of all-time downloads
proxectonos/Nos_StyleTTS2-Brais-GL
Nos_StyleTTS2-Brais-GL is a text-to-speech model from proxectonos. Use it when you need text read aloud. The card lists the license as apache-2.0.
- Model Description - System Architecture and Components - Intended Uses and Limitations - Prerequisites and Environment Setup - Auxiliary Components - Model Training
Downloads · 30 days
9
100% of all-time downloads
All-time downloads
9
Public
Repo size
8.7 GB
Likes
0
Public
Click a slice to open those files.
.pth3.9 GB · 79%
From the Hugging Face model README
Nos_StyleTTS2-Brais-GL is a high-quality text-to-speech (TTS) model for Galician based on the StyleTTS2 architecture. It was trained on the single-speaker male voice "Brais" (proxectonos/Nos_Brais-GL), developed as part of Proxecto Nós.
This model integrates a native Galician processing pipeline, utilizing Cotovía for grapheme-to-phoneme (G2P) conversion and PL-ModernBERT-gl as its phoneme-level language encoder, providing rich contextualized phoneme embeddings.
proxectonos/PL-ModernBERT-gl (ModernBERT trained on Galician phonemes).proxectonos/Nos_Brais-GL.The synthesis pipeline relies on four main modules:
proxectonos/PL-ModernBERT-gl trained with Cotovía's phoneme dictionary to supply contextual embeddings.BSC-LT/hubert-base-los-2k to calculate Speech Language Model losses during second-stage adversarial training.git clone [https://huggingface.co/proxectonos/Nos_StyleTTS2-Brais-GL](https://huggingface.co/proxectonos/Nos_StyleTTS2-Brais-GL)
cd Nos_StyleTTS2-Brais-GL
conda create -n StyleTTS2 python=3.10 -y
conda activate StyleTTS2
pip install -r requirements.txt
Utils/cotovia/README.md.Training data can be downloaded using the provided download script. You can specify the primary training dataset and the Out-Of-Domain (OOD) dataset in Configs/download_data.yml.
To download the dataset via the script:
python scripts/download_data.py --config Configs/download_data.yml --download_data
Note: If you wish to switch the primary dataset, keep the existing audio folders, delete the
.txtannotation files, and rerun the script without the--download_dataflag:
python scripts/download_data.py --config Configs/download_data.yml
To ensure optimal synthesis, StyleTTS2 relies on three auxiliary models aligned with the same phonemizer (Cotovía):
StyleTTS2 requires a PL-BERT model trained on Galician phonemes generated by Cotovía. The training code and utilities for this component are maintained in the proxectonos/PL-ModernBERT-gl repository.
An Auxiliary ASR model aligns phoneme sequences with acoustic features. Follow the instructions in Utils/ASR/AuxiliaryASR (a modified version of AuxiliaryASR adapted for Cotovía). You can reuse the dataset downloaded in the previous steps for retraining.
The Utils/JDC directory contains a pre-trained Joint Detection and Classification (JDC) model for pitch (F0) extraction. Although pre-trained on the English LibriTTS corpus, pitch dynamics are language-independent, making retraining unnecessary.
| Parameter | Value / Description |
|---|---|
first_stage_path | path_to_first_stage/model.pth |
F0_path | Utils/JDC/bst.t7 |
ASR_config | Utils/ASR/config.yml |
ASR_path | path_to_ASR/model.pth |
PLBERT_dir | Path to trained PL-BERT directory |
Important: For
PLBERT_dir, specify only the directory containing theconfig.ymlconfiguration file, which indicates the path to the BERT model.
| Parameter | Value | Parameter | Value |
|---|---|---|---|
epochs_1st | 100 | preprocess.sr | 24,000 Hz |
epochs_2nd | 100 | preprocess.n_fft | 2,048 |
batch_size | 4 | preprocess.win_length | 1,200 |
grad_accumulation | 1 | preprocess.hop_length | 300 |
max_len | 500 | model.n_mels | 80 |
load_only_params | false | model.n_token | 69 |
model.hidden_dim | 512 | model.style_dim | 128 |
model.multispeaker | false | model.dropout | 0.2 |
| Component | Parameter | Value |
|---|---|---|
| Decoder | type | istftnet |
resblock_kernel_sizes | [3, 7, 11] | |
upsample_rates | [10, 6] | |
upsample_initial_channel | 512 | |
| SLM Discriminator | model | BSC-LT/hubert-base-los-2k |
sr | 16,000 Hz | |
hidden | 768 | |
nlayers | 13 | |
| Diffusion | embedding_mask_proba | 0.1 |
transformer.num_layers | 3 | |
transformer.num_heads | 8 |
Training is executed in two consecutive stages:
accelerate launch train_first.py --config_path ./Configs/config.yml
python train_second.py --config_path ./Configs/config.yml
Note: Multi-GPU training is not supported in the second stage.
Checkpoints are saved in log_dir using the naming conventions epoch_1st_%05d.pth and epoch_2nd_%05d.pth.
This repository includes an inference script to synthesize audio using a pre-trained StyleTTS2-GL model. It accepts raw Galician text strings or input text files and supports configurable parameters to control the generation process.
python inference.py --config Configs/inference_config.yml --text "Texto en galego para sintetizar." --device 0
| Argument | Description | Default / Example |
|---|---|---|
--config | Path to the YAML inference configuration file | Configs/inference_config.yml |
--text | Text string to synthesize | "Texto a xerar" |
--file | Path to a file containing text samples to synthesize | path/to/file.txt |
--device | Target execution device (cpu or GPU index) | 0 |
--output_dir | Directory where output files will be saved | ./results |
--output_file | Filename for the generated audio file | output.wav |
--diffussion_steps | Number of diffusion sampling steps | 30 |
--embedding_scale | Scaling factor for the generated style vector | 1.0 |
--alpha | Weight factor for the reference style embedding | 0.3 |
--beta | Weight factor for the generated style embedding | 0.7 |
--t | Interpolation factor with the previous style embedding | 0.0 |
--evaluate | Enables automated MOS evaluation on generated audio | Flag |
The following parameters are recommended for optimal audio synthesis using the Brais StyleTTS2 model:
| Parameter | Argument | Recommended Value |
|---|---|---|
| $\alpha$ | --alpha | 0.6 |
| $\beta$ | --beta | 1.0 |
| $\theta$ | --t | 0.6 |
| Diffusion Steps | --diffusion_steps | 10 |
| Embedding Scale | --embedding_scale | 1.0 |
An evaluation script utilizing the speechmos library is included to assess audio quality and verify model performance via MOS (Mean Opinion Score).
To run the evaluation on generated audio files:
python scripts/eval.py --audios audio1.wav audio2.wav --output output/mos_results.txt
Depending on the number of audio files provided, the script performs the following calculations:
Results will be printed to the stdout console and saved in the designated --output file.
Results were obtained using the speechmos library to measure predicted DNSMOS (OVRL) quality and calculate CMOS gains.
Models were evaluated using three different text lengths containing a mix of neutral, interrogative, and exclamatory sentences:
Synthesized audio samples were benchmarked against original corpus recordings and baseline VITS models from Proxecto Nós.
| Dataset / Voice Corpus | DNSMOS Metric (OVRL) |
|---|---|
| Nos_Brais-GL (Original) | 3.308 |
| Text Length | VITS Model (MOS) | StyleTTS2 Model (MOS) | Net Gain (CMOS) |
|---|---|---|---|
| Short (~10s) | 3.275 | 3.426 | +0.151 |
| Medium (~30s) | 3.257 | 3.442 | +0.185 |
| Long (>60s) | 3.241 | 3.397 | +0.156 |
If this model contributes to your research, please cite it as follows:
@misc{proxectonos/Nos_StyleTTS2-Brais-GL,
author = {{Proxecto Nós}},
title = {{Nos_StyleTTS2-Brais-GL: Galician Male TTS Model}},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{[https://huggingface.co/proxectonos/Nos_StyleTTS2-Brais-GL](https://huggingface.co/proxectonos/Nos_StyleTTS2-Brais-GL)}},
}
This model is licensed under the Apache License 2.0.
This work is funded by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project Desarrollo de Modelos ALIA.
We would like to express our gratitude to the engineering and research teams at Gradiant for the technical development of this model, as well as to the Aholab Signal Processing Laboratory (HiTZ) and the Language Technologies Laboratory (BSC) for their technical support and collaboration.