Downloads · 30 days
43
100% of all-time downloads
VantoraLabs/Vocetta2-147k
Vocetta2-147k is a text-to-speech model from VantoraLabs. Use it when you need text read aloud. It is set up for pytorch. The card lists the license as mit.
A complete English text-to-speech system in 147,532 parameters — three tiny networks and a dictionary-based grapheme-to-phoneme front end, all running on CPU in real time.
Downloads · 30 days
43
100% of all-time downloads
All-time downloads
43
Public
Parameters
148K
5.7 MB on disk
Likes
0
Public
Click a slice to open those files.
.json6.3 MB · 52%
From the Hugging Face model README
A complete English text-to-speech system in 147,532 parameters — three tiny networks and a dictionary-based grapheme-to-phoneme front end, all running on CPU in real time.
| Params | 147,532 (duration 5,344 + acoustic 40,162 + decoder 102,026) |
| Audio | 24 kHz mono |
| Speed | ~134x real time on CPU (RTF 0.0075) |
| Intelligibility | WER 0.105 on a 24-sentence diverse held-out set (7/24 exact) |
| Naturalness | SCOREQ 1.48, DNSMOS-OVRL 2.95, DNSMOS-SIG 3.19 |
| License | MIT (runtime and weights); bundled G2P data is Apache-2.0 |
Listen to samples/ first — those eight files were rendered by this exact
checkpoint.
Text becomes audio through four stages, all trained by distillation from a larger teacher TTS:
text ──► G2P ──► phoneme ids
│
duration.pt │ ids ──────────────► frame count per phoneme
│
acoustic.pt │ ids + durations ───► mel spectrogram [100 bands, T frames]
│
decoder.pt │ mel + noise ───────► complex spectrum ──► iSTFT ──► audio
G2P (microtts/g2p/). A dictionary-first English front end: words are
looked up in bundled pronunciation dictionaries (gold and silver), and words
the dictionaries miss go through a small bundled neural fallback model. NumPy
only, no espeak, no torch, no network access. The output is a string of IPA
phonemes, mapped to a frozen 62-symbol vocabulary (<bos> and <eos> bracket
each utterance).
Duration student (duration.pt, 5,344 params). A 3-layer 1D convolutional
network over the phoneme sequence. It predicts how many mel frames each phoneme
occupies, using learned position, sequence-length and duration features, with
residual blocks around each conv pair. The output is exponentiated log-duration,
rounded and clamped to at least one frame per phoneme.
Acoustic student (acoustic.pt, 40,162 params). Embeds the phoneme ids,
refines them with token-context convolutions, expands them to the frame grid by
repeating each phoneme its predicted number of frames, then runs a second stack
of convolutions over frames and projects to 100 mel bands.
Decoder (decoder.pt, 102,026 params). Mel to waveform. A ConvNeXt1D stack
(depthwise conv, LayerNorm, two pointwise layers with a GELU between them,
residual) maps the mel to a complex spectrum of 513 bins, which the iSTFT turns
into audio. The magnitude head is exponential with bin 0 and the Nyquist bin
zeroed, and a DC-blocking filter removes the remaining offset. The decoder is
noise-fed: a 4-channel noise input is projected and added to the mel
embedding. At inference, zero noise is the best choice.
This model exists to answer one question: how much quality survives under 150,000 parameters?
Two allocations were measured at almost exactly the same total size. The one shipped here puts the money in the decoder:
| config | duration | acoustic | decoder | total | SCOREQ | WER |
|---|---|---|---|---|---|---|
| this model | 5,344 | 40,162 | 102,026 | 147,532 | 1.48 | 0.105 |
| alternative | 5,344 | 65,299 | 77,218 | 147,861 | 1.32 | 0.100 |
The wider decoder buys naturalness (+0.16 SCOREQ); the wider acoustic buys a little intelligibility (−0.005 WER). Decoder width dominated every capacity comparison made while building this family, which is why the budget went there. A ~276K model trained the same way reaches SCOREQ 2.04 — the extra ~130K is almost entirely decoder, and that is the honest cost of the remaining gap.
Every student is trained by distillation: a larger teacher TTS renders a text corpus once, and the students learn to reproduce the teacher's intermediate representations.
Pick a teacher TTS and a text corpus (thousands of sentences of varied,
spoken-style text). For each line store: phoneme ids, teacher audio, per-phoneme
durations, and the mel spectrogram of the teacher audio (100 bands, n_fft 1024,
hop 256). One .npz per line (train/build_pack.py). Watch the duration units
when the teacher's frame rate differs from the mel hop — the script shows the
conversion.
Train ids → frame counts against the teacher's durations. Loss: smooth-L1 on
log-duration plus a term on the total length (weight 0.35). Learning rate 2e-3
with AdamW, ~4k steps at batch 32. Capacity matters and not monotonically:
hidden 12 is the sweet spot this model uses; sweep a few sizes if you change the
corpus.
Train ids + durations → mel against the teacher's mel. Loss: L1 plus
normalized L1, temporal-delta L1 (weight 0.10 — this anti-smoothing term is
load-bearing), channel-statistics L1, and a hinge-loss PatchGAN critic on the
mel from step 1500 with weight 0.1. The learning rate matters most: 2e-3,
constant. ~30k steps at batch 8.
An extra step is worth naming: this acoustic was distilled against the mel output of a larger acoustic in the same family, not against the teacher's mel directly. Distilling against a reachable target (a function the student can actually represent) was worth more than any capacity change tried here.
train/init_decoder.py; dim 40, pw 120,
3 blocks for this model). Training a decoder this small from scratch does not
reach intelligibility.--mix-prob 0.0) with
waveform L1 + multi-resolution spectral loss + a hinge PatchGAN
discriminator + a cosine-gram temporal-structure loss (weight 0.4 — this
supervises temporal texture and is what keeps the output from sounding
mushy), plus high-band-excess and quiet-ceiling terms. ~20k steps at batch 4,
constant lr 2e-4.--mix-prob 0.5), another ~20k steps. Load-bearing for the handoff.Two traps measured during development, worth stating plainly:
Download this repository, then install the runtime dependencies and run the script below. Two commands do the download — pick whichever tool you have:
# either the Hugging Face CLI (install it if you don't have it)
pip install -U huggingface_hub
hf download VantoraLabs/Vocetta2-147k --local-dir vocetta2-147k
# or git-lfs (install git-lfs first: https://git-lfs.com)
git lfs install && git clone https://huggingface.co/VantoraLabs/Vocetta2-147k
cd vocetta2-147k
pip install -r requirements.txt # numpy, torch, soundfile
Then, from inside that folder:
from microtts import MicroTTS
tts = MicroTTS.load(".") # reads duration.pt / acoustic.pt / decoder.pt
wav = tts.synthesize("Hello world.") # float32 numpy array, 24 kHz
tts.save("out.wav", wav) # saves RMS-normalized to -26 dBFS
MicroTTS.load accepts device="cpu" (default) or "cuda". The full pipeline
loads in ~0.1 s and needs no network access. If you already have phoneme ids,
call tts.synthesize_ids(ids) directly and skip the G2P.
Two practical notes:
synthesize returns the raw output; its loudness is not
normalized. MicroTTS.normalize(wav) applies RMS normalization (0.05, about
-26 dBFS), and MicroTTS.save does it for you.synthesize(..., noise_scale=0.0) is the default and gives the
best measured quality. The decoder still accepts noise_scale=1.0 if you
want variation across renders, but it costs quality on every metric.Python 3.10+, CPU is enough.
Measured on this exact checkpoint. All word-error numbers use Whisper small as the judge; naturalness scores use the standard open models (SCOREQ, DNSMOS), each evaluated on the raw rendered audio.
Intelligibility (word error rate, lower is better)
| set | WER | sentences exactly right |
|---|---|---|
| 24-sentence diverse held-out set | 0.105 | 7 / 24 |
| 128-sentence templated eval set | 0.000 | 128 / 128 |
The templated set (near-identical sentence frames) is fully memorized; the diverse set is the honest generalisation measure and the number to compare against other models.
Naturalness / audio quality (higher is better, 24-sentence diverse set)
| setting | SCOREQ | DNSMOS-OVRL | DNSMOS-SIG |
|---|---|---|---|
| noise_scale = 0.0 (recommended) | 1.48 | 2.95 | 3.19 |
DNSMOS-SIG catches metallic distortion; 3.19 at zero noise says the output is not buzzy.
Speed (24 diverse sentences, warm-up excluded, zero noise, CPU)
| device | RTF | real-time factor |
|---|---|---|
| CPU | 0.0075 | ~134x faster than real time |
RTF includes the duration, acoustic and decoder forwards; the G2P adds ~1 ms per sentence on top.
Reproducing the scores. Word error: transcribe the rendered wavs with
openai/whisper-small and compute WER against the input text (scripts in
benchmark/). Naturalness: pip install scoreq speechmos, then score each wav
with the library's own defaults.
duration.pt, acoustic.pt, decoder.pt the weights (606 KB total, fp32)
model.safetensors same weights, prefixed keys, fp32
(auto-detected by HF Hub so the params
count shows on the repo card)
microtts/ runtime package (frontend, models, g2p)
g2p/g2p_data/ dictionaries + fallback model, Apache-2.0
samples/ eight rendered examples
train/ the training scripts (see the recipe above)
benchmark/ RTF + WER measurement scripts and results
README.md, LICENSE, requirements.txt
Runtime, weights and training scripts: MIT (this release). The bundled
grapheme-to-phoneme dictionaries and the fallback model are Apache-2.0; see
microtts/g2p/g2p_data/NOTICE.md.