Downloads · 30 days
0
Pendrokar/xvapitch
xvapitch is a text-to-speech model from Pendrokar. Use it when you need text read aloud.
GitHub project, inference Windows/Electron app: https://github.com/DanRuta/xVA-Synth
Downloads · 30 days
0
Access
Public
Updated May 13, 2025
Repo size
1.1 GB
Likes
2
Public
Click a slice to open those files.
.pt1.1 GB · 100%
From the Hugging Face model README
GitHub project, inference Windows/Electron app: https://github.com/DanRuta/xVA-Synth
Fine-tuning app: https://github.com/DanRuta/xva-trainer
The base model for training other 🤗 xVASynth's "xVAPitch" type models (v3). Model itself is used by the xVATrainer TTS model training app and not for inference. All created by Dan "@dr00392" Ruta.
The v3 model now uses a slightly custom tweaked VITS/YourTTS model. Tweaks including larger capacity, bigger lang embedding, custom symbol set (a custom spec of ARPAbet with some more phonemes to cover other languages), and I guess a different training script. - Dan Ruta
When used in xVASynth editor, it is an American Adult Male voice. Default pacing is too fast and has to be adjusted.
xVAPitch_5820651 model sample: <audio controls> <source src="https://huggingface.co/Pendrokar/xvapitch/resolve/main/xVAPitch_5820651.wav?download=true" type="audio/wav"> Your browser does not support the audio element. </audio>
There are hundreds of fine-tuned models on the web. But most of them use non-permissive datasets.
Papers:
Referenced papers within code:
Used datasets: Unknown/Non-permissiable data