Downloads · 30 days
14
16% of all-time downloads
TurkishCodeMan/xtts-v2-english-finetuned
xtts-v2-english-finetuned is a text-to-speech model from TurkishCodeMan. Use it when you need text read aloud. It is set up for coqui-tts. The card lists the license as apache-2.0.
This is a fine-tuned version of Coqui XTTS v2 for English text-to-speech synthesis.
Downloads · 30 days
14
16% of all-time downloads
All-time downloads
87
Public
Repo size
5.6 GB
Likes
1
Public
Click a slice to open those files.
.pth5.6 GB · 100%
From the Hugging Face model README
This is a fine-tuned version of Coqui XTTS v2 for English text-to-speech synthesis.
| Parameter | Value |
|---|---|
| Batch Size | 4 |
| Learning Rate | 5e-06 |
| Max Audio Length | 11 seconds |
| Total Training Samples | 168 |
| Epoch | Eval Loss |
|---|---|
| 0 | 3.36 |
| 1 | 3.23 |
| 2 | 3.17 |
| 3 | 3.12 |
| 4 | 3.10 |
| 5 | 3.08 |
| 6 | 3.07 |
| 7 | 3.07 (best) |
| 8 | 3.11 |
| 9 | 3.10 |
pip install TTS==0.22.0 torch==2.5.1 torchaudio==2.5.1 transformers==4.40.0
pip install huggingface_hub
import os
import torch
import torchaudio
from huggingface_hub import hf_hub_download
from TTS.tts.configs.xtts_config import XttsConfig
from TTS.tts.models.xtts import Xtts
# Download model files
repo_id = "TurkishCodeMan/xtts-v2-english-finetuned"
model_path = hf_hub_download(repo_id=repo_id, filename="model.pth")
config_path = hf_hub_download(repo_id=repo_id, filename="config.json")
vocab_path = hf_hub_download(repo_id=repo_id, filename="vocab.json")
# Load model
config = XttsConfig()
config.load_json(config_path)
model = Xtts.init_from_config(config)
model.load_checkpoint(
config,
checkpoint_dir=os.path.dirname(model_path),
checkpoint_path=model_path,
vocab_path=vocab_path,
use_deepspeed=False
)
model.cuda()
# Generate speech (download a sample reference audio first)
ref_audio = hf_hub_download(repo_id=repo_id, filename="samples/speaker_reference.wav")
gpt_cond_latent, speaker_embedding = model.get_conditioning_latents(audio_path=ref_audio)
out = model.inference(
text="Hello, this is a test of the fine-tuned XTTS model.",
language="en",
gpt_cond_latent=gpt_cond_latent,
speaker_embedding=speaker_embedding,
)
wav = torch.tensor(out["wav"]).unsqueeze(0)
torchaudio.save("output.wav", wav, 24000)
| Type | File |
|---|---|
| Speaker Reference | speaker_reference.wav |
| Generated Output | generated_output.wav |
⚠️ Important: Use specific versions to avoid compatibility issues.
trainer/generic_utils.py or use monkey-patch before importing TTS.CUDA_VISIBLE_DEVICES=0 before imports.Apache 2.0