Downloads · 30 days
120
18% of all-time downloads
shawnpi/HQ-SVC
HQ-SVC is a audio-to-audio model from shawnpi. Use it for the audio-to-audio task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Official Repository of Paper: "Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios"(AAAI 2026) <div align="center" <p <img src="images/kon-new.gif" alt="HQ-SVC Logo" width="300" </p <a hr…
Downloads · 30 days
120
18% of all-time downloads
All-time downloads
665
Public
Repo size
5.9 GB
Likes
10
Public
Click a slice to open those files.
.gz4.8 GB · 81%
From the Hugging Face model README
Official Repository of Paper: "Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios"(AAAI 2026)
<div align="center"> <p> <img src="images/kon-new.gif" alt="HQ-SVC Logo" width="300"> </p> <a href="https://arxiv.org/abs/2511.08496"><img src="https://img.shields.io/badge/arXiv-2511.08496-b31b1b.svg?logo=arxiv&logoColor=white" alt="arXiv"></a> <a href="https://shawnpi233.github.io/HQ-SVC-demo"><img src="https://img.shields.io/badge/Demos-🌐-blue" alt="Demos"></a> <a href="https://huggingface.co/shawnpi/HQ-SVC"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Models%20-%20Access-orange" alt="Models Access"></a> <a href="https://github.com/ShawnPi233/HQ-SVC" target="_blank" rel="noopener noreferrer"> <img src="https://img.shields.io/badge/GitHub-Repository-blue?logo=github" alt="GitHub Repository"></a> </div>HQ-SVC is an efficient framework for high-quality zero-shot singing voice conversion (SVC) in low-resource scenarios. It achieves disentanglement of content and speaker features via a unified decoupled codec, and enhances synthesis quality through multi-feature fusion and progressive optimization.
Unlike existing methods that demand large datasets or heavy computational resources, HQ-SVC unifies:
Tested only on Linux platforms with CUDA >= 11.8 (仅在 Linux 平台、CUDA >= 11.8 的环境上测试通过)
Windows users can use WSL (Ubuntu) for deployment and execution (Windows 用户可以使用 WSL (Ubuntu) 进行部署运行)
git clone https://github.com/ShawnPi233/HQ-SVC.git
cd HQ-SVC
wget -c https://huggingface.co/shawnpi/HQ-SVC/resolve/main/environment.tar.gz
wget -c https://hf-mirror.com/shawnpi/HQ-SVC/resolve/main/environment.tar.gz # Optional mirror
mkdir -p venv
tar -xzf environment.tar.gz -C venv
source venv/bin/activate
export HF_ENDPOINT=https://hf-mirror.com # Optional mirror
python gradio_app.py
Caught signal 11 (Segmentation fault: address not mapped to object at address (nil)) (如果报错 Caught signal 11 (Segmentation fault: address not mapped to object at address (nil)))unset LD_LIBRARY_PATH
<div align="center">
<img src="images/sr.png" alt="sr" width="500">
Zero-shot Super-Resolution (16 kHz to 44.1 kHz): Input only source audio
Zero-shot Singing Voice Conversion: Input both source audio and target audio
If you use HQ-SVC in your research, please cite our work:
@inproceedings{bai2026hqsvc,
author = {Bingsong Bai and Yizhong Geng and Fengping Wang and Cong Wang and Puyuan Guo and Yingming Gao and Ya Li},
title = {HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios},
booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
volume = {40},
number = {36},
pages = {30013--30021},
year = {2026},
doi = {10.1609/aaai.v40i36.40249},
url = {https://doi.org/10.1609/aaai.v40i36.40249}
}
We thank the open-source communities behind: