Downloads · 30 days
5
15% of all-time downloads
MLA299/Tennda-Waves
Tennda-Waves is a text-to-speech model from MLA299. Use it when you need text read aloud. It is set up for TTS. The card lists the license as other.
XTTS-v2 voice cloning model, fine-tuned via LoRA on coqui/XTTS-v2.
Downloads · 30 days
5
15% of all-time downloads
All-time downloads
33
Public
Repo size
2.1 GB
Likes
0
Public
Click a slice to open those files.
.pth2.1 GB · 100%
From the Hugging Face model README
XTTS-v2 voice cloning model, fine-tuned via LoRA on coqui/XTTS-v2.
from TTS.tts.models.xtts import Xtts
from TTS.tts.configs.xtts_config import XttsConfig
import torchaudio
# Load model
config = XttsConfig()
config.load_json("config.json")
model = Xtts.init_from_config(config)
model.load_checkpoint(
config,
checkpoint_path="model.pth",
vocab_path="vocab.json",
speaker_file_path=" ",
eval=True,
)
# Get speaker embeddings
gpt_cond_latent, speaker_embedding = model.get_conditioning_latents(
audio_path=["reference.wav"], # Reference audio
gpt_cond_len=30,
gpt_cond_chunk_len=4,
max_ref_length=60,
)
# Synthesize speech
out = model.inference(
text="Hello, this is Tennda speaking.",
language="en",
gpt_cond_latent=gpt_cond_latent,
speaker_embedding=speaker_embedding,
temperature=0.75,
)
# Save audio
torchaudio.save("output.wav", torch.tensor(out["wav"]).unsqueeze(0), 24000)
This model is fine-tuned from coqui/XTTS-v2, which is licensed under the CPML license (non-commercial research use only).
⚠️ Warning: This model is NOT licensed for commercial use.