Downloads ยท 30 days
0
ebellob/ZipVoice-CA
ZipVoice-CA is a text-to-speech model from ebellob. Use it when you need text read aloud. The card lists the license as apache-2.0.
Catalan fine-tune of ZipVoice, a fast zero-shot text-to-speech model based on flow matching.
Downloads ยท 30 days
0
Access
Public
Updated Apr 24, 2026
Repo size
491 MB
Likes
2
Public
Click a slice to open those files.
.pt491 MB ยท 100%
From the Hugging Face model README
Catalan fine-tune of ZipVoice, a fast zero-shot text-to-speech model based on flow matching.
<p align="center"> <a href="https://erikupv.github.io/zipvoice-samples/"> <img src="https://img.shields.io/badge/๐%20Listen-Samples-green" alt="Listen to samples"> </a> <a href="https://github.com/ErikUPV/ZipVoice-CA"> <img src="https://img.shields.io/badge/GitHub-ZipVoice--CA-orange?logo=github&logoColor=white" alt="GitHub repository"> </a> <a href="https://github.com/k2-fsa/ZipVoice"> <img src="https://img.shields.io/badge/Base%20Model-ZipVoice-blue" alt="Base ZipVoice repository"> </a> </p>This repository contains the fine-tuned ZipVoice-CA checkpoint for Catalan speech synthesis. For the full training, preprocessing, inference, and evaluation recipe, see the GitHub repository.
The metrics below are intended as indicative benchmarks under this repository's evaluation setup, not as definitive state-of-the-art claims.
| Dataset | WER (%) โ | CER (%) โ | SIM-o โ | UTMOS โ |
|---|---|---|---|---|
| Common Voice 17 | 10.96 | 3.00 | 0.68 | 3.17 |
| FestCat | 7.31 | 2.56 | 0.65 | 3.46 |
| LaFrescat | 7.61 | 2.56 | 0.67 | 3.54 |
Evaluation uses generated samples from the ZipVoice-CA recipe with guidance_scale=1.0 and num_step=25.
git clone https://github.com/ErikUPV/ZipVoice-CA.git
cd ZipVoice-CA
conda create -n ZipVoice python=3.11
conda activate ZipVoice
pip install -r requirements_zipvoice.txt
# pip install huggingface_hub
huggingface-cli download \
--local-dir models \
ebellob/ZipVoice-CA \
zipvoice_ca.pt
test.tsv filepython3 -m zipvoice.bin.infer_zipvoice \
--model-name zipvoice \
--model-dir ./models \
--checkpoint-name zipvoice_ca.pt \
--tokenizer espeak \
--lang ca \
--test-list data_cat/raw/test.tsv \
--res-dir results/ \
--guidance-scale 1.0 \
--num-step 25
python3 -m zipvoice.bin.infer_zipvoice \
--model-name zipvoice \
--prompt-wav prompt.wav \
--prompt-text "I am the transcription of the prompt wav." \
--text "I am the text to be synthesized." \
--res-wav-path result.wav \
--model-dir ./models \
--checkpoint-name zipvoice_ca.pt \
--tokenizer espeak \
--lang ca \
--guidance-scale 1.0 \
--num-step 25
The prompt audio should contain the reference speaker voice, and --prompt-text should match the transcription of that prompt audio.
The reported metrics are computed on generated samples from three Catalan evaluation sources:
Metrics:
This model is intended for Catalan text-to-speech research and experimentation. Quality may vary depending on prompt quality, prompt duration, speaker characteristics, text normalization, and out-of-domain inputs.
As with any zero-shot TTS model, users should avoid generating speech that impersonates real people without consent.
This model is a fine-tuned version of ZipVoice, using the pretrained checkpoint released by the original authors.
This model is released under the Apache-2.0 License.