Downloads · 30 days
0
HOLOGRAMTECH/q-pocket-tts
q-pocket-tts is a text-to-speech model from HOLOGRAMTECH. Use it when you need text read aloud. The card lists the license as cc-by-4.0.
TTS by Kyutai. A 100M-parameter CPU text-to-speech model, re-hosted as a content-addressed κ-object for the Hologram serverless substrate so the Bonsai demo can speak its answers as it writes them — 100% in the browse…
Downloads · 30 days
0
Access
Public
Updated Jul 16, 2026
Repo size
240 MB
Likes
1
Public
Click a slice to open those files.
.onnx125 MB · 51%
From the Hugging Face model README
TTS by Kyutai. A 100M-parameter CPU text-to-speech model, re-hosted as a content-addressed κ-object for the Hologram serverless substrate so the Bonsai demo can speak its answers as it writes them — 100% in the browser, on the CPU, leaving the GPU to the mind.
This is the int8 ONNX decode graph (community export by
KevinAHM) plus one voice (alba, Kyutai's
default), run with onnxruntime-web 1.20.0 (WASM/SIMD). No account, no data leaves the device;
warm after one visit.
Each file is gzipped and stored as a content-addressed block b/sha256_<hex>.gz; the manifest lists
them. Every block is verified against its content address before it is decompressed (Law L5).
| role | file | what it is |
|---|---|---|
flow_lm_main | flow_lm_main_int8.onnx | the streaming flow-LM backbone (autoregressive latents) |
flow_lm_flow | flow_lm_flow_int8.onnx | the stateless flow-matching step |
text_conditioner | text_conditioner_int8.onnx | text → conditioning embeddings |
mimi_decoder | mimi_decoder_int8.onnx | Mimi codec decoder → 24 kHz PCM |
tokenizer | tokenizer.model | SentencePiece tokenizer |
bos_before_voice | bos_before_voice.npy | BOS conditioning prefix |
bundle | bundle.json | runtime metadata (state manifests, sample rate) |
voice | voice.alba.bin | ONE voice state (PTVB1, single voice) — extracted from the 8-voice voices.bin |
manifest.json (sha256 transport axis) + manifest.blake3.json + sha256-to-blake3.map.json
(canonical BLAKE3 κ axis). Any static host can serve this — the bytes carry identity, not the host.
did:holo:sha256:fb8d982d3a37f845915358928f985b83339dd198b45b6f6a705bdc8145b1afd7sha256:a4c11cfbef09c5b30dc97733c11ec21e2573c3adc467fcf4409c06b44d6f41d4alba · language english_2026-04 · 24 kHz · latent dim 32131.7 MB raw → 75.7 MB stored (gzip), 8 blocks. Decode graph only — the voice-cloning encoder is omitted (a non-goal). One voice ships; swapping the voice is a one-file re-mint.
Chrome / WASM-SIMD, 24-core desktop: first sound 332 ms on a short clause single-threaded (1.07× realtime); 313 ms and 1.39× with threads (cross-origin isolation). Real speech, 24 kHz. Running concurrently with Bonsai-8B on WebGPU costs ~20% of decode throughput (CPU-orchestration, not GPU contention) — still real-time.
kyutai/pocket-tts (weights CC-BY-4.0), voices from
kyutai/tts-voices.KevinAHM/pocket-tts-onnx (export scripts Apache-2.0); upstream Pocket TTS
code is MIT.KevinAHM/pocket-tts-web demo space.holo-bonsai/mint-pocket-tts.mjs (gzip each opaque blob → sha256(gz) = block address;
root = sha256 over the sorted name:kappa set) + forge/gen-blake3-manifest.mjs (BLAKE3 axis).
No tensor re-layout — the ONNX graphs pass through byte-exact.See LICENSE (CC-BY-4.0) and NOTICE.txt.