Downloads · 30 days
0
akkiisfrommars/EvoTalk
EvoTalk is a machine learning model from akkiisfrommars. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A single-speaker text-to-speech acoustic model built from scratch in PyTorch. A standalone architecture that predicts mel spectrograms from phonemes, vocoded natively with Vocos at 24 kHz.
Downloads · 30 days
0
Access
Public
Updated Aug 14, 2026
Repo size
164 MB
Likes
0
Public
Click a slice to open those files.
.pt163 MB · 100%
From the Hugging Face model README
A single-speaker text-to-speech acoustic model built from scratch in PyTorch. A standalone architecture that predicts mel spectrograms from phonemes, vocoded natively with Vocos at 24 kHz.
The current v1 model has overfitted to its training data (Hi-Fi TTS speaker 9017), so generalization to new speakers or unseen prosody is limited. A second version addressing this is in active development. The model does however produce intelligible audio output for single-speaker synthesis.
CosmicFish-like: transformer encoder and decoder (GQA, RoPE, SwiGLU, RMSNorm) around a variance adaptor that predicts per-phoneme duration, pitch, and energy, then expands the enriched sequence to frame level with a length regulator. ~80M parameters.
| Stage | Script | Purpose |
|---|---|---|
| Prepare | prepare.py | G2P, alignment, feature extraction, dataset + metadata |
| Train | train.py | Main training with masked L1/MSE losses, AMP, cosine LR |
| Finetune | finetune.py / long_finetune.py | Long-form and extra-long utterance stages |
| Synthesize | inference.py | Interactive REPL with sentence/clause chunking |
| Predict | predict.py | Render the predicted mel spectrogram as an image |
| Sweep | sweep.py | Render one script across every checkpoint |
python prepare.py --out_dir data
python train.py --data_dir data
python inference.py
Install dependencies with pip install -r requirements.txt.
Training curves from graphs/.



Predicted mel spectrogram from python predict.py --text "Hello! How are you? I am EvoTalk a TTS model built by Mistyoz AI." --out mel.png.

Apache 2.0