Downloads · 30 days
0
ACloudCenter/F5-TTS-robot
F5-TTS-robot is a machine learning model from ACloudCenter. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Full finetune of F5TTSv1Base on ACloudCenter/robot-tts-corpus: ~1.7 h of synthetic robot speech in 6 styles (omnibot vocoder-buzz, terminator ring-mod android, scifi metallic, gruff, intercom PA, telephone narrowband)…
Downloads · 30 days
0
Access
Public
Updated Sep 22, 2026
Repo size
5.4 GB
Likes
0
Public
Click a slice to open those files.
.pt5.4 GB · 100%
From the Hugging Face model README
Full finetune of F5TTS_v1_Base on ACloudCenter/robot-tts-corpus: ~1.7 h of synthetic robot speech in 6 styles (omnibot vocoder-buzz, terminator ring-mod android, scifi metallic, gruff, intercom PA, telephone narrowband), built from 5 LibriSpeech test-clean speakers.
12 epochs (~3900 updates), lr 1e-5, batch 1618 frames, vocab = Emilia pinyin.
The checkpoint is a full merged model (no PEFT needed):
from f5_tts.api import F5TTS
model = F5TTS(ckpt_file="model_last.pt", vocab_file="vocab.txt")
# ref_audio: a clip in the damaged style you want (see ACloudCenter/robot-tts-corpus refs/)
wav, sr, _ = model.infer(ref_file="ref_omnibot.wav", ref_text="<its transcript>", gen_text="Attention. Reactor core temperature exceeds tolerance.")
Pick the reference clip matching the style you want generated - the style identity comes entirely from the reference audio. Held-out reference clips per style are in ACloudCenter/robot-tts-corpus (refs/ + refs.csv).