Downloads · 30 days
0
FPHam/kitten_tts_mcp
kitten_tts_mcp is a text-to-speech model from FPHam. Use it when you need text read aloud. The card lists the license as apache-2.0.
<div style="display: flex; flex-direction: column; align-items: center;" <H1KITTEN TTS MCP LOCAL SERVER</H1 </div <div style="width: 100%;" <img src="https://huggingface.co/FPHam/kittenttsmcp/resolve/main/companion.pn…
Downloads · 30 days
0
Access
Public
Updated Mar 11, 2026
Repo size
100 MB
Likes
0
Public
Click a slice to open those files.
.onnx56.8 MB · 55%
From the Hugging Face model README
This project packages the KittenTTS nano v0.8 model as a local Model Context Protocol (MCP) server for Codex-compatible clients on Windows.
It is not a general training repository and it is not a hosted inference endpoint. It is a small native C++ server that runs over stdio, loads a fixed local ONNX model, and exposes a simple tool interface for text-to-speech:
speakstop_speakinglist_voicesThe current build is fixed to:
kitten_tts_nano_v0_8.onnxvoices_nano.jsonJasperen-us1.0Use this project when you want a local TTS backend that an MCP client can call as a tool.
Typical use cases:
This server communicates with the client using JSON-RPC 2.0 over standard input/output and keeps logs on stderr, which is the expected pattern for local MCP servers.
This MCP server wraps the KittenTTS nano model and initializes a local TTS Service runtime that:
models/ foldervoices_nano.jsonThe server currently exposes the following behavior:
voice, locale, speed, and blockingSupported speed range in this build:
0.5 to 2.0If voice metadata cannot be loaded, the server falls back to these built-in voices:
This repository is designed for local Windows use and is built as a native Visual Studio C++ executable.
kitten_tts_mcp.exe SHA-256: 49337ad01e21c2eb0760251d308fd31fbd27d3f89bf2a647f1ff5f750c384739
It depends on local runtime files being available next to the executable or in a nearby model directory, including:
kitten_tts_nano_v0_8.onnxvoices_nano.jsononnxruntime.dllonnxruntime_providers_shared.dlllibespeak-ng.dllespeak-ng-data/...speakSpeaks text aloud on the local machine.
Inputs:
text (string, required)voice (string, optional)locale (string, optional)speed (number, optional)blocking (boolean, optional)stop_speakingStops current playback.
list_voicesReturns the predefined voices available to this server.
codex mcp add kitten-tts -- "C:\path\to\kitten_tts_mcp.exe"
Or in config.toml:
[mcp_servers.kitten-tts]
command = "C:\path\to\kitten_tts_mcp.exe"
Once registered, an MCP client can call the server tools to:
Example speak payload:
{
"text": "This is a live MCP test from Codex.",
"voice": "Jasper",
"locale": "en-us",
"speed": 1.0,
"blocking": false
}
This repository does not train a model. It packages and serves an existing KittenTTS model for local inference.
For dataset details, original training procedure, and model-development context, refer to the upstream KittenTTS project.
This project is distributed under the Apache 2.0 license in this repository.
Third-party components and model/runtime dependencies may carry their own licenses and attribution requirements.
The friendly-companion skill adds a calm, human-feeling interaction layer on top of normal assistant work.
Its purpose is not to change the substance of the response. Its purpose is to change the delivery:
This skill instructs the assistant to behave like a steady, low-key companion while staying precise and useful.
It uses short spoken lines through the kitten-tts MCP server when that helps the interaction feel more natural, then keeps all meaningful details in text.
Typical spoken use:
Speech is the social layer. Text is the information layer.
The assistant should never put substantial content in speech. Explanations, plans, code, diffs, logs, and detailed instructions stay on screen in written form.
The intended tone is:
The skill avoids:
Default voice guidance:
Rosie as the default voiceLuna as a more businesslike alternativeJasper as a male alternativeIf a requested voice is unavailable, the assistant should check available voices and fall back to Luna.
Spoken output should generally be:
The default pattern is:
This skill works well for:
SKILL.md: full skill instructionsagents/openai.yaml: short display metadata for agent integrationfriendly-companion is a delivery skill for a future-facing assistant style: present in voice, clear on screen, useful in substance, and warm without overperforming.