Downloads · 30 days
0
appautomaton/fireredtts3-mlx
fireredtts3-mlx is a text-to-speech model from appautomaton. Use it when you need text read aloud. It is set up for mlx. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Sep 27, 2026
Repo size
6.2 GB
Likes
0
Public
Click a slice to open those files.
.safetensors6.1 GB · 100%
From the Hugging Face model README
Multilingual voice cloning on Apple Silicon
Base BF16 · Pure MLX inference · Mono 24 kHz audio
Runtime guide · Source code · Upstream model · Project website
</div>This repository brings FireRedTTS3 Base to mlx-speech as a complete MLX voice-cloning pipeline. Give it a reference recording, its transcript, and the text you want spoken. It returns a waveform in the reference speaker's voice.
App Automaton maintains the MLX conversion and runtime. Speech generation runs locally on your Mac, including speaker conditioning and waveform reconstruction.
The current release is Base BF16, stored in base/mlx-bf16/. Each model
variant has its own complete inference bundle. Instruct will be added under
instruct/ after its MLX pipeline is validated, keeping the Base path stable.
Requires an Apple Silicon Mac and Python 3.13 or later.
pip install "mlx-speech>=0.5.3"
The loader downloads the Base bundle on first use. Replace reference.wav
and reference_text with your recording and its exact transcript. A reference
in the target language is preferred when available.
from mlx_speech import tts
from mlx_speech.audio import write_wav
model = tts.load(
"appautomaton/fireredtts3-mlx",
artifact_subdir="base/mlx-bf16",
)
result = model.generate(
"你好,很高兴认识你。",
reference_audio="reference.wav",
reference_text="For Timothy was a spoiled cat, and he allowed no one.",
language="Chinese",
seed=1234,
flow_steps=10,
guidance_scale=2.0,
)
write_wav("generated.wav", result.waveform, sample_rate=result.sample_rate)
The aliases fireredtts3-base and fireredtts3-base-bf16 select the same
bundle. Only base/mlx-bf16/ and the root model card are downloaded. Additional
variants in this repository do not increase the size of a Base download.
mlx-speech tts \
--model appautomaton/fireredtts3-mlx \
--artifact-subdir base/mlx-bf16 \
--text "你好,很高兴认识你。" \
--reference-audio reference.wav \
--reference-text "For Timothy was a spoiled cat, and he allowed no one." \
--language Chinese \
--seed 1234 \
--flow-steps 10 \
--guidance-scale 2.0 \
--output generated.wav
</details>
The model directory contains all three components and the tokenizer. No separate codec or speaker-model download is needed.
| Component | Weight file | Size |
|---|---|---|
| Qwen3 and DiT speech-generation core | core.safetensors | 3.950 GiB |
| RedAE waveform encoder and decoder | redae.safetensors | 1.758 GiB |
| CAM++ speaker encoder | speaker.safetensors | 0.014 GiB |
config.json, tokenizer.json, tokenizer_config.json, and vocab.json sit
beside the weight files at the directory root. The three weight files total
5.722 GiB. Inference also needs memory for activations and the KV cache.
README.md
LICENSE
base/
mlx-bf16/
config.json
core.safetensors
redae.safetensors
speaker.safetensors
tokenizer.json
tokenizer_config.json
vocab.json
A downloaded base/mlx-bf16/ directory also loads directly by local path.
Base is the only model variant included in this release.
The conversion casts the original FP32 trainable weights to BF16 and maps their names and layouts for MLX. RedAE's ISTFT window and CAM++ running statistics remain FP32. This artifact uses 16-bit floating-point weights and does not apply INT8 or INT4 quantization.
The tokenizer accepts the upstream model's 24 language identifiers and 21
Chinese dialect tags. Pass an explicit language such as "English",
"Chinese", or "Japanese" with each request.
Arabic, Cantonese, Chinese, Czech, Dutch, English, Finnish, French, German, Greek, Hindi, Indonesian, Italian, Japanese, Korean, Polish, Portuguese, Romanian, Russian, Spanish, Thai, Turkish, Ukrainian, and Vietnamese.
The dialect identifiers use the upstream ZH_ prefix, for example
"ZH_Sichuan" and "ZH_Shanghai". The complete list is defined in the
MLX tokenizer.
Language coverage comes from the upstream model and tokenizer. The local quality check below covers one Mandarin request with an English reference.
On the fixed request documented in the runtime guide:
| Check | Result |
|---|---|
| Output | Finite, non-silent, mono 24 kHz waveform |
| Seed repeatability | Two bitwise-identical waveforms on one loaded model |
| Local ASR transcript | 你好,很高兴认识你。 |
| CAM++ reference/output cosine | 0.7374 |
This covers one fixture only; it does not establish voice quality across other languages and recordings.
The conversion uses
FireRedTeam's weights at dcf1bdcd
and follows the
official implementation at 1d32ba78.
Both revisions are recorded in config.json.
The repository also retains small, deterministic golden fixtures for RedAE, CAM++, DiT, and autoregressive generation. They let maintainers check numerical regressions without keeping or downloading the full checkpoint. The fixture manifest records the BF16 bundle's SHA-256 hashes and capture provenance. Capture and replay use an explicit MLX CPU stream to keep these checks consistent between developer Macs and CI.
These fixtures exercise tiny models with reproducible synthetic weights. Checkpoint compatibility and full-model voice quality are covered by separate tests that require the real weights. The fixture guide documents the capture script and regression command.
FireRedTTS3 is developed by the FireRed Team and released under Apache 2.0. The MLX runtime and conversion code are maintained by App Automaton under the MIT license.
The upstream project describes voice cloning as intended for academic research. Use reference recordings with the speaker's permission and identify synthetic speech clearly. See the upstream model card for its intended-use statement.