Downloads · 30 days
0
Mearman/dac-encoder-onnx
dac-encoder-onnx is a machine learning model from Mearman. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
ONNX export of the encoder + quantizer from ibm-research/DAC.speech.v1.0.
Downloads · 30 days
0
Access
Public
Updated May 6, 2026
Repo size
882 KB
Likes
0
Public
Click a slice to open those files.
.onnx882 KB · 100%
From the Hugging Face model README
ONNX export of the encoder + quantizer from ibm-research/DAC.speech.v1.0.
Used for on-device voice cloning with OuteTTS 1.0. Encodes reference audio into discrete codebook indices that condition the OuteTTS model during synthesis.
weights_24khz_1.5kbps_v1.0.pth from ibm-research/DAC.speech.v1.0(batch, 1, samples) float32 PCM at 24kHz, values in [-1, 1](batch, 2, frames) int64 codes, values 0-1023import onnxruntime as ort
import numpy as np
sess = ort.InferenceSession("dac_encoder_24khz.onnx")
# 1 second of audio at 24kHz
audio = np.random.randn(1, 1, 24000).astype(np.float32) * 0.1
codes = sess.run(None, {"audio": audio})
# codes[0].shape = (1, 2, 75) — 2 codebooks, 75 frames
The codes map to special tokens in the OuteTTS 1.0 vocabulary:
<|c1_0|> through <|c1_1024|><|c2_0|> through <|c2_1024|>These are interleaved per frame to create the speaker conditioning prompt.
MIT (same as DAC)