Downloads · 30 days
1.9K
5% of all-time downloads
kenpath/svara-tts-voiceclone-beta
svara-tts-voiceclone-beta is a text-to-speech model from kenpath. Use it when you need text read aloud. It is set up for transformers. The card lists the license as apache-2.0.
[](https://huggingface.co/kenpath/svara-tts-voiceclone-beta) [](https://huggingface.co/spaces/kenpath/svara-tts) [](https://colab.research.google.com/)
Downloads · 30 days
1.9K
5% of all-time downloads
All-time downloads
41.4K
Public
Parameters
3.3B
6.6 GB on disk
Likes
11
Public
Click a slice to open those files.
.safetensors6.6 GB · 100%
From the Hugging Face model README
svara-tts-voiceclone-beta is an experimental extension of svara-tts-v1, designed to bring lightweight voice cloning and improved accent preservation to Indic languages. It introduces a simple but effective reference-swap finetuning technique, enabling more stable zero-shot speaker identity across long, expressive utterances.
Built on an Orpheus-style discrete audio token architecture, the model supports 19 languages, expressive cues (<laugh>, <yawn>, <angry>), and low-latency TTS on commodity hardware.
Demo playback uses the same Space as svara-tts-v1.
आज शाम को जल्दी मिलते हैं। <neutral>Zero-shot example:
<BOS>
<reference_audio_tokens_here>
कल शाम को जल्दी मिलते हैं। <neutral>
<SOA>
Speaker IDs remain compatible with svara-tts-v1: Language (Gender).
svara-tts-voiceclone-beta is enhanced from the multilingual base of svara-tts-v1, trained on:
The reference-swap augmentation uses multi-utterance samples to improve speaker consistency across Indic phonetic variation.
These improve with targeted LoRA finetuning or higher-quality data.
By using this model, you agree to follow applicable laws and ethical guidelines. Synthetic speech should be disclosed when appropriate. Avoid impersonation or harmful use cases.
Developed by Kenpath Technologies. Special thanks to:
Apache-2.0