Downloads · 30 days
0
FenixDS/styletts2-spanish-ft
styletts2-spanish-ft is a machine learning model from FenixDS. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Modelo StyleTTS2 fine-tuned para síntesis de voz en español con clonación de voz.
Downloads · 30 days
0
Access
Public
Updated Jan 8, 2026
Repo size
2.3 GB
Likes
0
Public
Click a slice to open those files.
.pth2.3 GB · 100%
From the Hugging Face model README
Modelo StyleTTS2 fine-tuned para síntesis de voz en español con clonación de voz.
Este modelo fue entrenado específicamente para generar voz en español con alta calidad y naturalidad. Incluye capacidades de clonación de voz mediante audio de referencia.
epoch_2nd_00049.pth: Checkpoint del modelo (2.1GB)config_spanish_ft.yml: Configuración del modeloreference_audio.wav: Audio de referencia para clonación de voz (916KB)pip install -U "huggingface_hub[cli]"
huggingface-cli download FenixDS/styletts2-spanish-ft --local-dir styletts2-spanish-ft
Este modelo está diseñado para usarse con VoxBridge. Configuración en config/default.yaml:
tts:
provider: styletts2
config_path: styletts2-spanish-ft/config_spanish_ft.yml
checkpoint_path: styletts2-spanish-ft/epoch_2nd_00049.pth
reference_audio: styletts2-spanish-ft/reference_audio.wav
alpha: 0.3
beta: 0.5
diffusion_steps: 4
embedding_scale: 2
import torch
from styletts2 import tts
# Cargar modelo
model = tts.StyleTTS2(
config_path="styletts2-spanish-ft/config_spanish_ft.yml",
checkpoint_path="styletts2-spanish-ft/epoch_2nd_00049.pth"
)
# Generar voz
text = "Hola, este es un ejemplo de síntesis de voz en español."
reference_audio = "styletts2-spanish-ft/reference_audio.wav"
audio = model.inference(
text=text,
ref_audio=reference_audio,
alpha=0.3,
beta=0.5,
diffusion_steps=4,
embedding_scale=2
)
MIT License
Basado en StyleTTS2 por yl4579.
Fine-tuning realizado con datos de voz en español guatemalteco.