Downloads · 30 days
0
OpenVoiceOS/phoonnx-kitten-tts
phoonnx-kitten-tts is a text-to-speech model from OpenVoiceOS. Use it when you need text read aloud. It is set up for kittentts. The card lists the license as apache-2.0.
phoonnx mirror of KittenML/KittenTTS — a family of tiny, fully offline, CPU-only English TTS models. The upstream ONNX graphs are used unmodified; this repo adds phoonnx-canonical config.json files (vocab + inference…
Downloads · 30 days
0
Access
Public
Updated Aug 4, 2026
Repo size
214 MB
Likes
1
Public
Click a slice to open those files.
.onnx214 MB · 100%
From the Hugging Face model README
phoonnx mirror of
KittenML/KittenTTS — a family of tiny,
fully offline, CPU-only English TTS models. The upstream ONNX graphs are used
unmodified; this repo adds phoonnx-canonical config.json files (vocab +
inference defaults) and splits each voices.npz into one style .bin file
per voice, matching phoonnx's StyleTTS2Adapter engine_params contract.
Upstream sources (Apache-2.0, verified on each model card):
nano-0.1/model.onnx, config.json, <voice>.bin (x8)
nano-0.2/model.onnx, config.json, <voice>.bin (x8)
mini-0.1/model.onnx, config.json, <voice>.bin (x8)
Voices (same 8 across all tiers): expr-voice-2-m, expr-voice-2-f,
expr-voice-3-m, expr-voice-3-f, expr-voice-4-m, expr-voice-4-f,
expr-voice-5-m, expr-voice-5-f.
input_ids (int64) + style (1, 256) + speed (1) -> waveform @ 24kHz, the same
single-graph contract phoonnx's StyleTTS2Adapter already implements for
StyleTTS2 and Kokoro — no new adapter code was needed. Tokenization is a
per-character map over espeak-ng en-us IPA (with stress marks), reproduced
verbatim from kittentts==0.1.3's inline vocab, not hand-copied.
WER (parakeet-tdt-0.6b-v2, 5 sentences, expr-voice-2-f):
| tier | WER | mean RTF (CPU) |
|---|---|---|
| nano-0.1 | 0.08 | 0.24 |
| nano-0.2 | 0.08 | 0.23 |
| mini-0.1 | 0.04 | 0.49 |
No torch reference implementation is published upstream (KittenTTS ships ONNX only), so there is no torch<->onnxruntime parity check to run. Upstream ships exactly one precision per tier (no separate fp16/int8 variants), so no variant-consistency check applies either.
No 5000-sample trailing trim is applied: kittentts==0.1.3's own
KittenTTS.generate() returns the raw graph output with no post-processing,
and inspecting the tail of phoonnx-synthesized audio shows no noise burst
(tail RMS in the same range as the rest of the clip).