Downloads · 30 days
0
BSC-LT/PL-BERT-ca
PL-BERT-ca is a machine learning model from BSC-LT. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<details <summaryClick to expand</summary
Downloads · 30 days
0
Access
Public
Updated Mar 6, 2026
Repo size
11.6 GB
Likes
0
Public
Click a slice to open those files.
.t710.5 GB · 100%
From the Hugging Face model README
PL-BERT-ca is a phoneme-level masked language model trained on Catalan text with diverse regional accents. It is based on the PL-BERT architecture, which learns phoneme representations via a BERT-style masked language modeling objective.
This model is designed to support phoneme-based text-to-speech (TTS) systems, including but not limited to StyleTTS2. Thanks to its Catalan-specific phoneme vocabulary and contextual embedding capabilities, it can serve as a phoneme encoder for any TTS architecture requiring phoneme-level features.
Features of our PL-BERT:
token_maps.pkl and adapted util.pyHere is an example of how to use this model within the StyleTTS2 framework:
Clone the StyleTTS2 repository: https://github.com/yl4579/StyleTTS2
Inside the Utils directory, create a new folder, for example: PLBERT_cat_multiaccent
Copy the following files into that folder:
config.yml (training configuration)step_1000000.t7 (trained checkpoint)token_maps.pkl (phoneme to ID mapping)util.py (modified to fix position ID loading)In your StyleTTS2 configuration file, update the PLBERT_dir entry to:
PLBERT_dir: Utils/PLBERT_cat_multiaccent
Update the import statement in your code to:
from Utils.PLBERT_cat_multiaccent.util import load_plbert
Phonemize your Catalan text files for training and validation (if you use espeak-ng use the language code ca)
Note: Although this example uses StyleTTS2, the model is compatible with other TTS architectures that operate on phoneme sequences. You can use the contextualized phoneme embeddings from PL-BERT in any compatible synthesis system.
The model was trained on a phonemized Catalan corpus (any phonemizer can be used) extracted from the CATalog corpus. The dataset includes sentences from speakers across Catalonia, Balearic Islands, and Valencia. It uses a consistent phoneme token set with boundary markers and masking tokens.
Tokenizer: custom (split using whitespaces)
Phoneme masking strategy: word-level and phoneme-level masking and replacement
Training steps: 1,000,000
Precision: Mixed (fp16)
Model parameters:
Other parameters:
The model has not been benchmarked via perplexity or extrinsic evaluation, but has been successfully integrated into TTS pipelines such as StyleTTS2, where it enables the synthesis of Catalan with regional accent variation.
If this code contributes to your research, please cite the work:
@misc{zevallos2025plbertca,
title={PL-BERT-ca},
author={Rodolfo Zevallos, Jose Giraldo and Carme Armentano-Oller},
organization={Barcelona Supercomputing Center},
url={https://huggingface.co/langtech-veu/PL-BERT-ca},
year={2025}
}
The Language Technologies Laboratory of the Barcelona Supercomputing Center by Rodolfo Zevallos.
For further information, please send an email to [email protected].
Copyright(c) 2025 by Language Technologies Laboratory, Barcelona Supercomputing Center.
This work is funded by the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project Desarrollo de Modelos ALIA.