Downloads · 30 days
0
Lyodos/classic-vc
classic-vc is a machine learning model from Lyodos. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
ClassicVC is an any-to-any voice conversion model that enables users to design their original speaker styles by selecting the coordinates from the continuous latent spaces. The model components are implemented using P…
Downloads · 30 days
0
Access
Public
Updated Aug 5, 2024
Repo size
1.4 GB
Likes
0
Public
Click a slice to open those files.
.onnx934 MB · 68%
From the Hugging Face model README
ClassicVC is an any-to-any voice conversion model that enables users to design their original speaker styles by selecting the coordinates from the continuous latent spaces. The model components are implemented using PyTorch and fully compatible with ONNX.
MMCXLI provides the dedicated graphical user interface (GUI) for ClassicVC. It runs on wxPython and ONNX Runtime. Users can download the ONNX files and try out speech conversion without having to install PyTorch or train a model with their own voice data.
Based on the MIT License, users can use the model codes and checkpoints for research purpose. It is provided with no guarantees.
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->This model was prototyped as a hobbyist's research into any-to-any voice conversion, and we make no guarantees especially regarding its reliability or real-time operation.
As for use in situations involving an unspecified number of people, such as web broadcasting, and mission-critical applications, including medical, transportation, infrastructure, and weapon systems, we cannot prohibit such use as the developer, since the MIT License is the only stated license, but we do not encourage it.
[More Information Needed]
We used three large-scale speech corpora (LibriSpeech, Samrómur Children 21.09, and VoxCeleb 1 and 2) to make the latent space of speakers that can be embedded using the style encoder of ClassicVC as inclusive as possible of all natural human voice.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
The Notebook 01 of the ClassicVC repository provides the procedure for offline (non real-time) voice conversion.
The MMCXLI repository provides GUI, which depends on local Python environment.
The model checkpoints provided here were trained on the following three datasets.
The Notebook 02 of the ClassicVC repository provides the procedure for data preparation.
The Notebook 03 of the ClassicVC repository provides the training code.