Downloads · 30 days
4.8K
32% of all-time downloads
cstr/voxcpm2-GGUF
voxcpm2-GGUF is a text-to-speech model from cstr. Use it when you need text read aloud. The card lists the license as apache-2.0.
GGUF conversion of openbmb/VoxCPM2 for use with CrispASR.
Downloads · 30 days
4.8K
32% of all-time downloads
All-time downloads
15.1K
Public
Repo size
12.5 GB
Likes
8
Public
Click a slice to open those files.
.gguf12.5 GB · 100%
From the Hugging Face model README
GGUF conversion of openbmb/VoxCPM2 for use with CrispASR.
# Zero-shot TTS
crispasr -m voxcpm2-f16.gguf \
--tts "Hello, this is VoxCPM2 speaking." \
--tts-output output.wav
# Quantized (smaller, faster)
crispasr -m voxcpm2-q4_k.gguf \
--tts "Hello world" --tts-output output.wav
| File | Size | Description |
|---|---|---|
voxcpm2-f16.gguf | 4.63 GB | F16 weights (full precision) |
voxcpm2-q4_k.gguf | ~1.5 GB | Q4_K quantized (faster, slightly lower quality) |
voxcpm2-q8_0.gguf | 2.83 GB | Q8_0 quantized (near-F16 quality) |
voxcpm2-q8_0-locdit-f16.gguf | 3.03 GB | Q8_0 with the LocDiT diffusion head kept in F16 — experimental mixed-precision Vulkan recipe (see below) |
voxcpm2-ref.gguf | 371 KB | Reference activation dump for numerical validation |
Historical T4 Vulkan component measurements found CFM per audio step of 70.2 ms
for Q8_0 versus 59.4 ms for the mixed recipe at ten steps. An eight-step run
measured 46.7 ms. These component timings do not establish speech quality or
an accepted end-to-end speedup. The mixed file was built with
crispasr-quantize voxcpm2-f16.gguf out.gguf q8_0 --tensor-type "^locdit\.=f16".
A previous CPU synthesis roundtrip is not GPU acceptance.
Validation update, 2026-10-09: matched real T4 tests at ten steps and seed 2 produced repeated tokens in the Nemotron ASR transcripts of the short test sentence for both Q8_0 and the mixed recipe. Six of sixteen exact primary ASR cases passed. This rejected gate is not proof that synthesized audio audibly stutters: a later fixed-audio comparison found both Parakeet and Qwen3 transcribe every one of 22 original Python/native CPU/GPU controls exactly across three resampling paths, while Nemotron retains errors. Resampling alone does not resolve the discrepancy. The completed original NVIDIA float32 and native F16 comparison reads all 22 VoxCPM2 controls correctly; native Q4 reads only 8/22. Original and F16 normalized words agree on all 26 tested clips, and fresh/reused sessions agree. The repeated tokens are therefore attributable to the native Nemotron Q4 recognizer on these fixed controls. Critical RNNT precision guards are being tested before changing that quant or its defaults. The published Q4 primary gate remains rejected while that recognizer is repaired. This diagnosis does not certify eight steps, Intel B390 performance, or complete original/native waveform parity. Defaults remain ten steps.
Original NVIDIA/F16/Q4 inputs, features, encoder states and complete transcripts
(SHA256 7762f2a73d8ae7dad05d667fbf0bca393a0b3b7ad622b1b1f09473fa8f5d0463).
All 234 fixed-audio transcripts, recognizer pins and resampling controls
(SHA256 3dd2f988e327dd30bf3ae608edc7611e52aedf3a8374a3ea6e2f288342110b98).
Immutable raw PCM, transcripts, backend traces and recipe evidence
(SHA256 f52e5fbc49a0654bdcd13da6660c8e0841fbe87e80eb0cc158d74801339a5abc).
See CrispASR #461 for hardware follow-up.
The voxcpm2-ref.gguf file contains intermediate activation tensors captured from the PyTorch reference implementation. Used with crispasr-diff to validate the C++ inference path:
crispasr-diff voxcpm2-tts voxcpm2-f16.gguf voxcpm2-ref.gguf samples/jfk.wav
Historical F16 transformer-stage check: 12 pass, 0 fail (TSLM, RALM, LocEnc, LocDiT, projections). This stage check does not certify complete generated speech, the current mixed quant, or every GPU backend.
Captured stages: text_input_ids, locenc_in, locenc_out, enc_to_lm, tslm_layer_0_out, tslm_layer_27_out, tslm_prefill_out, ralm_prefill_out, lm_to_dit_hidden, res_to_dit_hidden, cfm_step0_z, cfm_step0_result, stop_logits_step0.
Text → BPE tokenize → TSLM (28L causal MiniCPM-4, GQA 16h/2kv, LongRoPE)
↓
FSQ bottleneck (tanh→round→linear)
↓
RALM (8L causal, no RoPE, GQA 16h/2kv)
↓
Projections: lm_to_dit + res_to_dit → mu [2048]
↓
LocDiT (12L bidirectional, CFM Euler solver, 10 steps, cfg=2.0)
↓
Predicted latent patch [4 frames × 64 dims]
↓
LocEnc (12L bidirectional) → next TSLM input
↓ (AR loop until stop)
AudioVAE decoder → 48 kHz PCM
Converted using models/convert-voxcpm2-to-gguf.py from the CrispASR repository:
python models/convert-voxcpm2-to-gguf.py \
--input openbmb/VoxCPM2 \
--output voxcpm2-f16.gguf
Original model by OpenBMB. Apache 2.0 license.
openbmb.apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not.