Downloads · 30 days
2
4% of all-time downloads
lyrin/voxpm2-ncnn
voxpm2-ncnn is a text-to-speech model from lyrin. Use it when you need text read aloud. It is set up for ncnn. The card lists the license as apache-2.0.
This repository contains an NCNN export of openbmb/VoxCPM2 for use with the voxcpm2-ncnn C++ runtime.
Downloads · 30 days
2
4% of all-time downloads
All-time downloads
47
Public
Repo size
13.6 GB
Likes
0
Public
Click a slice to open those files.
.bin5.2 GB · 100%
From the Hugging Face model README
This repository contains an NCNN export of openbmb/VoxCPM2 for use with the voxcpm2-ncnn C++ runtime.
It is a converted runtime asset package, not a newly trained or fine-tuned model. The model weights keep the same Apache-2.0 license as the upstream VoxCPM2 release.
Model repository: https://huggingface.co/lyrin/voxpm2-ncnn
Runtime source: https://github.com/LiYulin-s/voxcpm2-ncnn.git
openbmb/VoxCPM2.param / .bin component graphsvoxcpm2-ncnnVoxCPM2 is a multilingual controllable speech generation model. The upstream release describes support for 30 languages, 9 Chinese dialects, voice design, style-controllable voice cloning, and high-fidelity continuation cloning. This NCNN package targets the modes currently exposed by the voxcpm2-ncnn runtime.
The exported model directory contains the runtime assets and a local LICENSE copy. The repository root keeps an additional LICENSE copy for hosting tools that expect the license at the top level.
model.json: NCNN component manifest and runtime settings*.ncnn.param, *.ncnn.bin: exported NCNN component graphs and weightstokenizer.json: tokenizer asset used by the runtimeLICENSE: Apache-2.0 license text for the model assetsExport-time intermediate files such as TorchScript, PNNX graphs, and generated Python wrappers are intentionally not included.
Download the model assets into the runtime repository:
huggingface-cli download lyrin/voxpm2-ncnn --local-dir assets/voxcpm2
Smoke-test the NCNN components:
xmake run voxcpm2 -m assets/voxcpm2 --smoke-components
xmake run voxcpm2 -m assets/voxcpm2 --smoke-components --vulkan
Generate speech from text:
xmake run voxcpm2 -m assets/voxcpm2 \
-t "你好,欢迎使用 VoxCPM2 NCNN。" \
-o out.wav
Use prompt continuation with prompt audio:
xmake run voxcpm2 -m assets/voxcpm2 \
-t "这是续写测试。" \
--prompt "你好。" \
--prompt-audio prompt.wav \
-o out.wav
Use reference audio:
xmake run voxcpm2 -m assets/voxcpm2 \
-t "这是参考音频测试。" \
--reference-audio reference.wav \
-o out.flac
The output format is inferred from the -o extension.
This package splits VoxCPM2 into NCNN component graphs:
The runtime uses a page-style KV cache internally while adapting to the current exported decoder-step NCNN graphs.
The original VoxCPM2 model is by OpenBMB / ModelBest. Please refer to the upstream VoxCPM2 model card, project repository, and technical report for model architecture, training, evaluation, and intended-use details.
The model assets in this repository are released under Apache-2.0, matching the upstream VoxCPM2 release. The license text is included in LICENSE.