Downloads · 30 days
15
42% of all-time downloads
walston/accentbox-zh
accentbox-zh is a text-to-speech model from walston. Use it when you need text read aloud. It is set up for tts. The card lists the license as apache-2.0.
Chinese AccentBox checkpoint for multi-accent TTS with independent speaker and accent references.
Downloads · 30 days
15
42% of all-time downloads
All-time downloads
36
Public
Repo size
1.1 GB
Likes
0
Public
Click a slice to open those files.
.pth1 GB · 96%
From the Hugging Face model README
Chinese AccentBox checkpoint for multi-accent TTS with independent speaker and accent references.
waveform_decoder). There is no separate external vocoder checkpoint.speaker_reference.walston/GenAID extracts a 64-d accent embedding from accent_reference.Main files:
best_model.pth: AccentBox acoustic model plus built-in HiFiGAN waveform decoder.config.json: training config.speaker_encoder/model_se.pth.tar: speaker encoder checkpoint.speaker_encoder/config_se.json: speaker encoder config.remote_infer.py: batch inference from a Joycent-style jsonl list.Use JSONL. Each row must contain:
{
"id": "case_id",
"speaker_reference": "/path/to/speaker.wav",
"accent_reference": "/path/to/accent.wav",
"phones": "sil n i3 h ao3 m a5 sil"
}
The Joycent inference list at resources/multi_accent/inference/medium_grl/infer_list.jsonl already has this format.
Install this AccentBox codebase and its Python dependencies, then run:
python remote_infer.py \
--repo-id walston/accentbox-zh \
--infer-list /path/to/infer_list.jsonl \
--output-dir outputs \
--device cuda
For CPU smoke tests:
python remote_infer.py \
--repo-id walston/accentbox-zh \
--infer-list /path/to/infer_list.jsonl \
--output-dir outputs \
--device cpu \
--limit 1
The script downloads this model repo, loads walston/GenAID, reads each list row, and writes one waveform per row to --output-dir.