Downloads · 30 days
5.8K
38% of all-time downloads
aufklarer/Chatterbox-Flash-CoreML
Chatterbox-Flash-CoreML is a text-to-speech model from aufklarer. Use it when you need text read aloud. It is set up for coremltools. The card lists the license as mit.
Compiled Core ML export of ResembleAI/chatterbox-flash for Apple runtimes.
Downloads · 30 days
5.8K
38% of all-time downloads
All-time downloads
15.3K
Public
Repo size
2.4 GB
Likes
0
Public
Click a slice to open those files.
.bin2.4 GB · 100%
From the Hugging Face model README
Compiled Core ML export of ResembleAI/chatterbox-flash for Apple runtimes.
This bundle contains the Chatterbox Flash T3 token generator and the S3Gen audio back-half as static-shape .mlmodelc graphs. It is intended for an application runtime that owns text tokenization, T3 denoising/sampling, reference conditioning, and graph orchestration.
| Component | Parameters | Format | Precision | Static shape |
|---|---|---|---|---|
| T3 block-diffusion token generator | 532.4M | Core ML .mlmodelc | fp16 | text_len 256, block_size 16, max_seq 1024 |
| S3Gen audio back-half | 266.0M | Core ML .mlmodelc | fp16 | token_len 192, mel_len 384 |
| Total | 798.4M | Core ML .mlmodelc | fp16 | 24 kHz waveform output |
| Path | Size | Description |
|---|---|---|
config.json | 4 KB | Root metadata for download tracking and runtime discovery |
t3/ConditioningEncoder.mlmodelc | 25 MB | Speaker/prompt/emotion conditioning to T3 conditioning embedding |
t3/TextPrefill.mlmodelc | 963 MB | Causal [cond, text, start_speech] prefix prefill and flat KV cache |
t3/BlockDecoder.mlmodelc | 1.0 GB | Full-visible Flash speech-block logits with explicit KV cache |
t3/uncond_block_prior.npy | 36 KB | Unconditional PMI prior for Flash scoring |
t3/tokenizer.json | 28 KB | Chatterbox Flash text tokenizer |
t3/config.json | 4 KB | T3 export metadata |
audio/FlowSpeakerProjector.mlmodelc | 44 KB | S3Gen reference embedding to projected speaker conditioning |
audio/FlowEncoder.mlmodelc | 79 MB | prompt_token ++ speech_tokens to flow mu and mask |
audio/FlowEstimator.mlmodelc | 141 MB | One meanflow Euler derivative step |
audio/HiFTVocoder.mlmodelc | 41 MB | Mel frames to 24 kHz waveform |
audio/audio_config.json | 4 KB | S3Gen audio export metadata |
The exported graphs cover:
The host runtime must still provide:
ref.wav -> speaker_embref.wav -> prompt_speech_tokensref.wav -> prompt_token, prompt_feat, embedding(t,r) = (0,0.5), (0.5,1.0)This means the bundle supports voice-cloning TTS when the runtime supplies reference conditioning tensors, but it is not yet a fully Core ML ref.wav -> cloned wav pipeline.
| Test | Result |
|---|---|
| T3 graph roundtrip vs PyTorch wrappers | Pass at 2% relative tolerance |
| Audio graph roundtrip vs PyTorch wrappers | Pass for token_len 192, mel_len 384 |
| Stitched S3Gen meanflow-to-audio roundtrip | Pass |
| Prompted synthesis smoke test | Pass |
| Whisper tiny transcript of generated wav | Core ML speech test. |
| Smoke-test WER | 0.000 |
Core ML warnings from local export:
CPU_ONLY prediction crashed for the T3 package in local coremltools 8.3 testing. Use ALL, CPU_AND_NE, or a compiled-device runtime..mlmodelc folders. The numerical parity tests were run against the source .mlpackage exports before compilation.The runtime loads graphs from t3/ and audio/, then:
ref_dict tensors from a prompt wav.t3/tokenizer.json.prompt_token and generated speech tokens.audio/FlowEncoder.mlmodelc.cond by copying prompt_feat.T into the mel prefix.audio/FlowEstimator.mlmodelc twice for (0,0.5) and (0.5,1.0).audio/HiFTVocoder.mlmodelc and crop padded samples.Converted from ResembleAI/chatterbox-flash, revision 4385507288b8197e6dab8b4e6b1603328d549d9d.