Downloads · 30 days
15
100% of all-time downloads
jjiang4/thorsten_vits
thorsten_vits is a text-to-speech model from jjiang4. Use it when you need text read aloud. It is set up for espnet. The card lists the license as cc-by-4.0.
VITS trained on the Thorsten-Voice Dataset 2022.10 German single male speaker corpus, using the egs2/thorsten/tts1 recipe in espnet. End-to-end, so no separate vocoder is needed.
Downloads · 30 days
15
100% of all-time downloads
All-time downloads
15
Public
Repo size
373 MB
Likes
0
Public
Click a slice to open those files.
.pth373 MB · 100%
From the Hugging Face model README
jjiang4/thorsten_vitsVITS trained on the Thorsten-Voice Dataset
2022.10 German single male speaker corpus,
using the egs2/thorsten/tts1 recipe in
espnet. End-to-end, so no separate vocoder
is needed.
Trained for 40k steps, 4% of the 1M-step schedule used for the LJSpeech VITS
model. Acoustic quality is already past the Tacotron 2 system at this point, but
intelligibility is not: alignment is learned by monotonic alignment search, which
converges much more slowly than the adversarial waveform objective. For the
better word error rate, use
jjiang4/thorsten_tts_train_tacotron2_raw_phn_espeak_ng_german
with jjiang4/thorsten_hifigan_ft_ljspeech.
This model was trained on pre-phonemised text (g2p: none), so it expects a
space-separated phoneme string, not raw German. Convert the text with the same
espeak-ng German frontend the recipe uses:
from espnet2.bin.tts_inference import Text2Speech
from espnet2.text.phoneme_tokenizer import PhonemeTokenizer
g2p = PhonemeTokenizer("espeak_ng_german")
tts = Text2Speech.from_pretrained(
"jjiang4/thorsten_vits",
noise_scale=0.333, # defaults (0.667, 0.8) are worse at this step count
noise_scale_dur=0.0,
)
text = "im prozess wurden aber nur vierzig fälle thematisiert."
wav = tts(" ".join(g2p.text2tokens(text)))["wav"]
Passing raw text runs but produces garbage: every word maps to the unknown token, and the utterance above comes out 1.0 s long instead of 3.2 s.
100-utterance test set. See the recipe README.
| System | MCD | log-F0 RMSE | UTMOS | WER (%) | CER (%) |
|---|---|---|---|---|---|
| Ground truth | - | - | 3.29 ± 0.21 | 6.2 | 3.2 |
| Tacotron 2 + HiFi-GAN | 5.83 ± 1.36 | 0.268 ± 0.061 | 3.04 ± 0.28 | 10.2 | 4.1 |
| VITS (40k steps) | 6.50 ± 0.92 | 0.255 ± 0.064 | 3.35 ± 0.23 | 16.2 | 6.5 |
espnet 202604pytorch 2.8.0+cu1283.10.14