Downloads · 30 days
7
35% of all-time downloads
OpenVoiceOS/phoonnx-neutts
phoonnx-neutts is a text-to-speech model from OpenVoiceOS. Use it when you need text read aloud. It is set up for phoonnx. The card lists the license as other.
ONNX conversion of afrispeech/Akiti-TTS, an Asante Twi text-to-speech model, packaged for phoonnx. This repository holds converted weights only — no new training was done.
Downloads · 30 days
7
35% of all-time downloads
All-time downloads
20
Public
Repo size
1.6 GB
Likes
0
Public
Click a slice to open those files.
.data977 MB · 63%
From the Hugging Face model README
ONNX conversion of afrispeech/Akiti-TTS, an Asante Twi text-to-speech model, packaged for phoonnx. This repository holds converted weights only — no new training was done.
Akiti-TTS is a LoRA fine-tune of pnnbao-ump/VieNeu-TTS-0.3B, which comes from Neuphonic's NeuTTS Air family: a Qwen3 causal LM that emits NeuCodec audio tokens.
akiti-twi-onnx/)| File | What it is |
|---|---|
neutts_lm.onnx + neutts_lm.onnx.data | Qwen3 backbone, fp32, KV-cached. The graph needs its .data sidecar next to it. |
neutts_lm_int8.onnx | The same graph, dynamically quantized to int8. Smaller, but measurably lower quality — it puts the end-of-speech token in the top 5 at the first step, where fp32 does not. |
neucodec_decoder.onnx | Copy of neuphonic/neucodec-onnx-decoder-int8 (Apache-2.0), unmodified. |
tokenizer.json | The checkpoint's own BPE, copied from upstream. |
voices.json | The nine voice presets, copied from AfriSpeech/akiti-tts (MIT). |
neutts_onnx_meta.json | Architecture summary written by the export script. |
neutts_lm.onnx serves prefill and decode: the same graph with a different past length.
inputs input_ids int64 [1, S] prompt tokens, or 1 token per step
attention_mask int64 [1, P + S] ones over past and current tokens
position_ids int64 [1, S] absolute positions, P .. P+S-1
past_key_<i> fp32 [1, 4, P, 64] i in 0..27
past_value_<i> fp32 [1, 4, P, 64]
outputs logits fp32 [1, 66938] last position only
present_key_<i> / present_value_<i> fp32 [1, 4, P + S, 64]
neucodec_decoder.onnx takes codes int32 [1, 1, N] and returns audio float32
[1, 1, 480 * (N - 1)] at 24 kHz — 50 codec tokens per second of audio.
The LM is prompted with phonemes, not letters. Text is phonemized by espeak-ng using the
lfn (Lingua Franca Nova) voice, which is what the checkpoint was trained with;
espeak-ng has no Twi voice, and lfn's five-vowel orthography reads Twi spelling closely.
<|TEXT_PROMPT_START|>{reference phones} {target phones}<|TEXT_PROMPT_END|>
<|SPEECH_GENERATION_START|>{<|speech_c|> for c in reference codes}
Generation continues that speech-token run until <|SPEECH_GENERATION_END|> or EOS.
Exported with scripts/conversion/neutts/export_neutts_onnx.py in phoonnx. Against the
torch model on a fixed prompt, max absolute logit difference:
| fp32 ONNX | |
|---|---|
| prefill | 4.10e-05 |
| decode (8 steps) | 2.96e-05 |
The fp32 graph's top-5 next tokens on a real prompt are identical to torch's, in the same order and to two decimal places.
The pieces carry different licenses, and one of them is inconsistent upstream. Read this before using the model.
afrispeech/Akiti-TTS weights — the Hugging Face model card declares
CC BY-NC 4.0 (non-commercial). The
GitHub README instead states the weights are
Apache-2.0. These two statements disagree. This mirror records both and resolves
neither; treat the stricter of the two (non-commercial) as binding until AfriSpeech
clarifies. The GitHub repository's code is MIT, which is not in dispute.neucodec_decoder.onnx — Apache-2.0, from Neuphonic.voices.json — from the MIT-licensed AfriSpeech/akiti-tts repository.Attribution: AfriSpeech / Ghana NLP (Akiti-TTS), pnnbao-ump (VieNeu-TTS-0.3B), Neuphonic (NeuTTS Air, NeuCodec).