Downloads · 30 days
0
walston/joycent-medium-grl
joycent-medium-grl is a text-to-speech model from walston. Use it when you need text read aloud. It is set up for pytorch. The card lists the license as mit.
This is a Joycent Mandarin accent TTS acoustic model trained using accent embeddings extracted by walston/whisaid-medium-grl. The released checkpoint is epoch 100.
Downloads · 30 days
0
Access
Public
Updated Aug 16, 2026
Repo size
228 MB
Likes
0
Public
Click a slice to open those files.
.pt228 MB · 100%
From the Hugging Face model README
This is a Joycent Mandarin accent TTS acoustic model trained using accent
embeddings extracted by
walston/whisaid-medium-grl.
The released checkpoint is epoch 100.
from huggingface_hub import hf_hub_download
checkpoint_path = hf_hub_download(
repo_id="walston/joycent-medium-grl",
filename="grad_100.pt",
)
Pass the downloaded checkpoint to joycent/inference_joycent.py with the
--acoustic-checkpoint argument. Full synthesis also requires the Joycent
vocoder and reference-audio feature extraction dependencies described in the
Joycent repository.
@misc{wang2026joycentdiffusionbasedaccenttts,
title={Joycent: Diffusion-based Accent TTS without Accented Phone Prediction},
author={Xintong Wang and Ye Wang},
year={2026},
eprint={2606.16417},
archivePrefix={arXiv},
primaryClass={cs.SD},
}