Downloads · 30 days
1.1K
100% of all-time downloads
cstr/supertonic-3-GGUF
supertonic-3-GGUF is a text-to-speech model from cstr. Use it when you need text read aloud. It is set up for supertonic. The card lists the license as openrail.
GGUF conversion of Supertone/supertonic-3 for the native ggml runtime in CrispASR (issue 434). Non-autoregressive flow-matching TTS, 44.1 kHz, 31 languages, ~99 M parameters, ten preset voices.
Downloads · 30 days
1.1K
100% of all-time downloads
All-time downloads
1.1K
Public
Repo size
200 MB
Likes
0
Public
Click a slice to open those files.
.gguf200 MB · 100%
From the Hugging Face model README
GGUF conversion of Supertone/supertonic-3 for the native ggml runtime in CrispASR (issue #434). Non-autoregressive flow-matching TTS, 44.1 kHz, 31 languages, ~99 M parameters, ten preset voices.
ONE file: the four networks (duration predictor, text encoder, vector estimator, vocoder), the unicode indexer, the NFKD text-normalisation tables and all ten voice styles (F1–F5, M1–M5) are embedded.
crispasr --backend supertonic -m supertonic3-f16.gguf \
--tts "Hello from Supertonic." -o out.wav
# language / voice / speed / flow steps
crispasr --backend supertonic -m auto --tts "Guten Tag." -l de \
--voice F2 --tts-speed 1.1 --tts-steps 8 -o out_de.wav
The model weights are OpenRAIL-M, inherited from Supertone/supertonic-3 (Supertone, Inc.). The license permits commercial use but carries USE RESTRICTIONS (see the base repo's LICENSE) — you are responsible for complying with them. The upstream sample code that defines the inference pipeline is MIT (supertone-inc/supertonic). This conversion changes the storage format only; all credit for the model belongs to Supertone.
Converted and validated on Kaggle (chr1str/crispasr-supertonic-434):
per-stage parity against the upstream onnxruntime pipeline (text encoder,
all 8 CFG flow steps and the vocoder), plus a TTS→ASR roundtrip with an
upstream-reference control arm. See validation/results-434.json.