Downloads · 30 days
0
notmax123/BlueV3-onnx
BlueV3-onnx is a text-to-speech model from notmax123. Use it when you need text read aloud. It is set up for onnx. The card lists the license as mit.
ONNX export of BlueV3 TTS for CPU / CUDA / TensorRT inference. Includes the vocoder (codec decoder).
Downloads · 30 days
0
Access
Public
Updated Jul 11, 2026
Repo size
390 MB
Likes
1
Public
Click a slice to open those files.
.onnx390 MB · 100%
From the Hugging Face model README
ONNX export of BlueV3 TTS for CPU / CUDA / TensorRT inference. Includes the vocoder (codec decoder).
TTS version: v1.7.3 · Sample rate: 44.1 kHz · Exported from PyTorch ckpt_step_767000 + AE ae_541000
| File | Role |
|---|---|
text_encoder.onnx | Phoneme IDs + style → text embedding |
vector_estimator.onnx | Flow-matching Euler step (CFG baked in) |
vocoder.onnx | Latent → 44.1 kHz waveform (codec) |
duration_predictor.onnx | Text + style_dp → duration (seconds) |
stats.npz | Latent mean / std / normalizer_scale |
uncond.npz | Unconditional tokens (CFG / debugging) |
tts.json | Runtime config |
PyTorch weights (no codec): notmax123/BlueV3
hf download notmax123/BlueV3-onnx --local-dir ./onnx_models
style_ttl [1, 50, 256] and style_dp [1, 8, 16] (from a style JSON / reference encoder).text_ids, text_mask.duration_predictor → seconds; divide by speed; convert to latent length with base_chunk_size=512, chunk_compress_factor=6.text_encoder(text_ids, style_ttl, text_mask) → text_emb.vector_estimator for N steps (e.g. 8). Output is the next latent state (CFG is inside the graph; do not apply CFG again).stats.npz:
z = (x / normalizer_scale) * std + mean # raw 144-d
optionally drop the last compressed frame, then vocoder(latent=z) → wav_tts.vocoder.onnx expects raw (unnormalized) 144-channel latents (normalizer_scale=1 inside the export).
# ONNX (ORT CUDA / CPU)
uv run python run_onnx_inference.py --onnx_dir ./onnx_models --speaker netsiga --steps 8
# TensorRT (after building engines from these ONNX files)
uv run python create_tensorrt.py --onnx_dir ./onnx_models --engine_dir trt_engines
uv run python benchmark_trt.py --style_json voice_styles/Rotem.json --steps 8 --out out.wav
text_encoder: text_ids, style_ttl, text_mask → text_emb
vector_estimator: noisy_latent, text_emb, style_ttl, latent_mask, text_mask, current_step, total_step → denoised_latent
duration_predictor: text_ids, style_dp, text_mask → duration
vocoder: latent [B, 144, T] → wav_tts
MIT (see frontmatter).