Downloads · 30 days
40
20% of all-time downloads
mlx-community/LongCat-AudioDiT-1B-4bit
LongCat-AudioDiT-1B-4bit is a text-to-speech model from mlx-community. Use it when you need text read aloud. It is set up for mlx-audio.
This model was converted to MLX format from meituan-longcat/LongCat-AudioDiT-1B using mlx-audio version 0.4.3.
Downloads · 30 days
40
20% of all-time downloads
All-time downloads
198
Public
Parameters
1.4B
1.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.4 GB · 100%
How the weights are stored.
U321.3B · 89%
From the Hugging Face model README
This model was converted to MLX format from meituan-longcat/LongCat-AudioDiT-1B using mlx-audio version 0.4.3.
Refer to the original model card for more details on the model.
pip install -U mlx-audio
from mlx_audio.tts.utils import load
model = load("mlx-community/LongCat-AudioDiT-1B-4bit")
result = next(model.generate("Hello, this is a test of AudioDiT."))
audio = result.audio # mlx array, 24kHz
Play audio directly:
from mlx_audio.tts.audio_player import AudioPlayer
player = AudioPlayer(sample_rate=24000)
result = next(model.generate("The quick brown fox jumps over the lazy dog."))
player.queue_audio(result.audio)
player.wait_for_drain()
player.stop()
Clone any voice using a reference audio sample and its transcript. Use guidance_method="apg" for best voice cloning quality:
result = next(model.generate(
text="Today is warm turning to rain, with good air quality.",
ref_audio="reference.wav",
ref_text="Transcript of the reference audio.",
guidance_method="apg",
cfg_strength=4.0,
steps=16,
))
result = next(model.generate(
text="今天晴暖转阴雨,空气质量优至良,空气相对湿度较低。",
steps=16,
cfg_strength=4.0,
))
| Parameter | Default | Description |
|---|---|---|
steps | 16 | Euler ODE solver steps. Higher = better quality, slower |
cfg_strength | 4.0 | Classifier-free guidance strength |
guidance_method | "cfg" | "cfg" for TTS, "apg" for voice cloning |
seed | 1024 | Random seed for reproducibility |
ref_audio | None | Reference audio for voice cloning (24kHz) |
ref_text | None | Transcript of the reference audio |
# Zero-shot TTS
python -m mlx_audio.tts.generate \
--model mlx-community/LongCat-AudioDiT-1B-4bit \
--text "Hello, this is a test of AudioDiT." \
--play
# Voice cloning
python -m mlx_audio.tts.generate \
--model mlx-community/LongCat-AudioDiT-1B-4bit \
--text "Today is warm turning to rain." \
--ref_audio reference.wav \
--ref_text "Transcript of the reference audio." \
--play
LongCat-AudioDiT weights and code are released under the MIT License.