Downloads · 30 days
4.5K
25% of all-time downloads
aufklarer/Supertonic-3-CoreML
Supertonic-3-CoreML is a text-to-speech model from aufklarer. Use it when you need text read aloud. It is set up for coreml. The card lists the license as openrail.
First-party CoreML export of Supertonic-3's four non-autoregressive flow-matching graphs, for on-device iOS / Apple Neural Engine. Built by our own pipeline (speech-models/stmodels): weights lifted from the Supertone/…
Downloads · 30 days
4.5K
25% of all-time downloads
All-time downloads
18K
Public
Repo size
449 MB
Likes
0
Public
Click a slice to open those files.
.bin396 MB · 99%
From the Hugging Face model README
First-party CoreML export of Supertonic-3's four non-autoregressive flow-matching graphs, for
on-device iOS / Apple Neural Engine. Built by our own pipeline (speech-models/stmodels): weights
lifted from the Supertone/supertonic-3 ONNX
initializers → PyTorch nn.Module → coremltools (mlprogram, FP32, iOS18+).
| Module | mlpackage | parity max|Δ| |
|---|---|---|
| Duration predictor | DurationPredictor.mlpackage | 7.2e-06 ✓ |
| Vector estimator (ODE denoiser) | VectorEstimator.mlpackage | 2.5e-03 ✓ |
| Vocoder | Vocoder.mlpackage | 3.0e-04 ✓ |
| Text encoder | TextEncoder.mlpackage | mean 2.5e-04 (max 2.5e-2 at isolated positions) |
Text/duration use fixed T=128 (relpos attention has T-dependent pad widths — pad/segment text to
128); vocoder + vector-estimator use a dynamic latent-length RangeDim. The host runs the flow-matching
ODE loop (vector_estimator ×total_steps) — the graphs contain no control flow. Assets to drive them:
tts.json, unicode_indexer.json (G2P-free tokenizer table), voice_styles/*.json.
FP32 = parity reference. For ANE residency, use the mixed-precision
Supertonic-3-CoreML-FP16— vocoder + duration FP16, text-encoder + vector-estimator FP32; measured transparent at 47–51 dB mag-STFT SNR.
Supertone/supertonic-3
(commit 3cadd1ee6394adea1bd021217a0e650ede09a323), Supertone Inc., arXiv:2503.23108 — OpenRAIL-M
(use-based restrictions carry over: no non-consensual impersonation/deepfakes, etc.).TTSInterface CoreML model.