Downloads · 30 days
14
25% of all-time downloads
idle-intelligence/pocket-tts-int8
pocket-tts-int8 is a text-to-speech model from idle-intelligence. Use it when you need text read aloud. It is set up for candle. The card lists the license as cc-by-4.0.
INT8 channel-wise quantized version of kyutai/pocket-tts-without-voice-cloning for browser-based TTS inference via WebAssembly.
Downloads · 30 days
14
25% of all-time downloads
All-time downloads
56
Public
Parameters
118M
139 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors139 MB · 100%
How the weights are stored.
I897.1M · 82%
From the Hugging Face model README
INT8 channel-wise quantized version of kyutai/pocket-tts-without-voice-cloning for browser-based TTS inference via WebAssembly.
| Original | INT8 | |
|---|---|---|
| File | tts_b6369a24.safetensors | model.safetensors |
| Size | 225 MB | 132 MB |
| Dtype | BF16 | I8 + BF16 scales |
| Reduction | — | 41% |
Per-output-channel INT8 quantization with BF16 scale factors.
Each weight tensor foo is split into two tensors:
foo (I8) — quantized weight valuesfoo_scale (BF16) — one scale factor per output channelDequantization: weight_bf16 = weight_i8 * scale_bf16
See quantize_config.json for machine-readable metadata.
| File | Size | Description |
|---|---|---|
model.safetensors | 132 MB | Quantized model weights |
tokenizer.model | 58 KB | SentencePiece unigram tokenizer |
quantize_config.json | <1 KB | Quantization parameters |
Voice embeddings are unchanged — use them from the original repo.
This model is designed for use with tts-web, a browser-based TTS engine built with Candle (Rust ML framework) and WebAssembly. Dequantization happens at load time in memory.
import safetensors
from safetensors.numpy import load_file
import numpy as np
tensors = load_file("model.safetensors")
for name in list(tensors.keys()):
if name.endswith("_scale"):
continue
scale_name = f"{name}_scale"
if scale_name in tensors:
weight_i8 = tensors[name].astype(np.float32)
scale = tensors[scale_name].astype(np.float32)
tensors[name] = (weight_i8 * scale).astype(np.float16)
del tensors[scale_name]
Based on Kyutai's Pocket TTS — a 100M parameter text-to-speech model.
This is an independent port by idle intelligence, not affiliated with or endorsed by Kyutai Labs.
CC-BY-4.0 (same as the original model).