Downloads · 30 days
147
34% of all-time downloads
chenxie95/X-VC
X-VC is a machine learning model from chenxie95. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
[](https://arxiv.org/abs/2604.12456) [](https://github.com/Jerrister/X-VC) [](https://x-vc.github.io)
Downloads · 30 days
147
34% of all-time downloads
All-time downloads
433
Public
Repo size
5 GB
Likes
7
Public
Click a slice to open those files.
.pt5 GB · 100%
From the Hugging Face model README
Official code release for X-VC: Zero-shot Streaming Voice Conversion in Codec Space.
git clone https://github.com/Jerrister/X-VC.git
cd X-VC
conda create -n xvc python=3.10 -y
conda activate xvc
pip install -U pip
pip install -r requirements.txt
Prepare:
Then set paths in configs/xvc.yaml, especially:
model.generator.semantic_encoder.encoder.from_pretrainedmodel.generator.semantic_encoder.cfgmodel.generator.speaker_encoder.pretrained_dirPut checkpoints under ckpts/, for example:
ckpts/
xvc.pt
bash scripts/infer_single.sh
Key arguments in this script:
current=0 for offline inference.current>0 for streaming inference.chunk/current/future/smooth control streaming behavior.Outputs are saved under save_dir (default: outputs/xvc_single).
Use scripts/batch_infer_seedtts_offline.sh.
bash scripts/batch_infer_seedtts_offline.sh
This script reports:
saved_dirtotal_rtfUse scripts/batch_infer_seedtts_stream.sh.
bash scripts/batch_infer_seedtts_stream.sh
This script reports:
saved_diravg_latency_msBefore training, prepare the required pretrained dependencies:
Then set corresponding paths in configs/xvc.yaml, especially:
model.generator.checkpointmodel.discriminator.checkpointOrganize your training/validation data in JSONL format and set:
datasets.traindatasets.valin configs/xvc.yaml.
You can adjust training behavior in:
configs/xvc.yaml (main training config)configs/ds_stage2.json (DeepSpeed config)Use scripts/train.sh.
bash scripts/train.sh
Notes:
configs/ds_stage2.json).configs/xvc.yaml.WANDB_API_KEY in scripts/train.sh before running if you use wandb logging.Training config points to JSONL files in configs/xvc.yaml:
datasets.traindatasets.valEach JSONL line should be a JSON object.
Required fields:
target_uttsource_wav_pathtarget_wav_pathOptional field:
source_uttMinimal example:
{"source_utt":"utt_0001","source_wav_path":"<path_to_source>","target_utt":"utt_0002","target_wav_path":"<path_to_target>"}
This codebase builds upon open-source components from SAC and the broader audio generation ecosystem.
If you find our work useful in your research, please consider citing:
@misc{zheng2026xvczeroshotstreamingvoice,
title={X-VC: Zero-shot Streaming Voice Conversion in Codec Space},
author={Qixi Zheng and Yuxiang Zhao and Tianrui Wang and Wenxi Chen and Kele Xu and Yikang Li and Qinyuan Chen and Xipeng Qiu and Kai Yu and Xie Chen},
year={2026},
eprint={2604.12456},
archivePrefix={arXiv},
primaryClass={eess.AS},
url={https://arxiv.org/abs/2604.12456},
}
This project is licensed under the MIT License.