Downloads · 30 days
0
Prasanna05/nano-codec-decoder-onnx
nano-codec-decoder-onnx is a text-to-speech model from Prasanna05. Use it when you need text read aloud. The card lists the license as apache-2.0.
ONNX-optimized decoder for the NeMo NanoCodec audio codec.
Downloads · 30 days
0
Access
Public
Updated Dec 31, 2025
Repo size
128 MB
Likes
0
Public
Click a slice to open those files.
.onnx128 MB · 100%
From the Hugging Face model README
ONNX-optimized decoder for the NeMo NanoCodec audio codec.
This model provides 2.5x faster inference compared to the PyTorch version for KaniTTS and similar TTS systems.
| Configuration | Decode Time/Frame | Speedup |
|---|---|---|
| PyTorch + GPU | ~92 ms | Baseline |
| ONNX + GPU | ~35 ms | 2.6x faster ✨ |
| ONNX + CPU | ~60-80 ms | 1.2x faster |
Real-Time Factor (RTF): 0.44x on GPU (generates audio faster than playback!)
pip install onnxruntime-gpu numpy
For CPU-only:
pip install onnxruntime numpy
import numpy as np
import onnxruntime as ort
# Load model
session = ort.InferenceSession(
"nano_codec_decoder.onnx",
providers=["CUDAExecutionProvider", "CPUExecutionProvider"]
)
# Prepare input
tokens = np.random.randint(0, 500, (1, 4, 10), dtype=np.int64) # [batch, codebooks, frames]
tokens_len = np.array([10], dtype=np.int64)
# Run inference
outputs = session.run(
None,
{"tokens": tokens, "tokens_len": tokens_len}
)
audio, audio_len = outputs
print(f"Generated audio: {audio.shape}") # [1, 17640] samples
from onnx_decoder_optimized import ONNXKaniTTSDecoderOptimized
# Initialize decoder
decoder = ONNXKaniTTSDecoderOptimized(
onnx_model_path="nano_codec_decoder.onnx",
device="cuda"
)
# Decode frame (4 codec tokens)
codes = [100, 200, 300, 400]
audio = decoder.decode_frame(codes) # Returns int16 numpy array
The decoder consists of two stages:
Dequantization (FSQ): Converts token indices to latent representation
Audio Decoder (HiFiGAN): Generates audio from latents
pip install onnxruntime-gpupip install onnxruntimetokens (int64): Codec token indices
[batch_size, 4, num_frames][0, 499] (FSQ codebook indices)tokens_len (int64): Number of frames
[batch_size]audio (float32): Generated audio waveform
[batch_size, num_samples][-1.0, 1.0]audio_len (int64): Audio length
[batch_size]Compared to PyTorch reference implementation:
Audio quality is virtually identical to PyTorch version.
Apache 2.0 (same as source model)