Skip to content
stt-tts-service logo

STT-TTS Service

stt-tts-service

Lightweight local speech-to-text and text-to-speech service for OpenClaw

SKILL.md

Full skill instructions

STT-TTS Service

A lightweight, local speech-to-text (STT) and text-to-speech (TTS) service that runs on any device connected to your OpenClaw server. Perfect for voice-enabled workflows and flexible resource allocation.

Features

  • Speech-to-Text: Transcribe audio using faster-whisper (4x faster than OpenAI Whisper)
  • Text-to-Speech: Generate natural speech using piper-tts or pyttsx3 fallback
  • 100% Local: No cloud APIs, works offline after initial model download
  • Flexible Deployment: Run on any device - Raspberry Pi, laptop, or GPU server
  • HTTP API: Simple REST endpoints for easy integration

Quick Start

Installation

# Clone or download this skill
cd stt-tts-service

# Install dependencies
pip install -r requirements.txt

# Start the service
python main.py

Docker Deployment

docker build -t stt-tts-service .
docker run -p 8765:8765 stt-tts-service

API Endpoints

POST /​stt - Speech to Text

Transcribe audio files to text.

curl -X POST http://localhost:8765/​stt \
  -F "[email protected]"

Response:

{
  "text": "Hello, this is the transcribed text.",
  "language": "en",
  "duration": 3.5
}

POST /​tts - Text to Speech

Convert text to audio.

curl -X POST http://localhost:8765/​tts \
  -H "Content-Type: application/​json" \
  -d '{"text": "Hello world", "voice": "default"}' \
  --output speech.wav

Parameters:

  • text (required): Text to synthesize
  • voice (optional): Voice ID to use
  • speed (optional): Speech rate multiplier (0.5-2.0)

GET /​health

Health check endpoint.

curl http://localhost:8765/​health

GET /​models

List available models and voices.

curl http://localhost:8765/​models

WebSocket Streaming (Real-time Voice)

For real-time voice conversations, use WebSocket endpoints:

WS /​ws/​stt - Streaming Speech-to-Text

Stream audio and receive transcriptions in real-time.

const ws = new WebSocket('ws://localhost:8765/​ws/​stt');

// Send audio chunks (16kHz, 16-bit, mono PCM)
ws.send(audioBuffer);

// Receive transcriptions
ws.onmessage = (event) => {
  const data = JSON.parse(event.data);
  console.log(data.text);  // Transcribed text
};

// Flush remaining audio
ws.send(JSON.stringify({action: "flush"}));

WS /​ws/​tts - Streaming Text-to-Speech

Send text and receive audio chunks in real-time.

const ws = new WebSocket('ws://localhost:8765/​ws/​tts');

// Send text to synthesize
ws.send(JSON.stringify({text: "Hello world"}));

// Receive audio chunks
ws.onmessage = (event) => {
  if (event.data instanceof Blob) {
    // Audio chunk - play it
    playAudio(event.data);
  }
};

WS /​ws/​voice - Full Duplex Voice Conversation

Stream audio input and receive audio output for real-time voice-to-voice.

const ws = new WebSocket('ws://localhost:8765/​ws/​voice');

// Stream microphone audio
navigator.mediaDevices.getUserMedia({audio: true})
  .then(stream => {
    // Send audio chunks to WebSocket
  });

// Handle responses
ws.onmessage = (event) => {
  const data = JSON.parse(event.data);
  if (data.type === "transcript") {
    // User's speech transcribed - send to your AI
    sendToAI(data.text);
  }
};

// Send AI response to be spoken
ws.send(JSON.stringify({action: "speak", text: aiResponse}));

Configuration

Set environment variables or edit config.py:

VariableDefaultDescription
STT_MODELbaseWhisper model: tiny, base, small, medium
TTS_ENGINEautoTTS engine: piper, pyttsx3, auto
DEVICEautoCompute device: cpu, cuda, auto
HOST0.0.0.0Server bind address
PORT8765Server port

Model Sizes

STT ModelSizeSpeedAccuracy
tiny~75MBFastestBasic
base~150MBFastGood
small~500MBMediumBetter
medium~1.5GBSlowerBest

OpenClaw Integration

Register this service with your OpenClaw server:

openclaw service register http://device-ip:8765

Then use in your workflows:

- action: stt
  input: ${audio_file}
  output: transcription
  
- action: tts
  input: "Hello, ${user_name}!"
  output: greeting_audio

Requirements

  • Python 3.9+
  • 2GB RAM minimum (4GB recommended for medium model)
  • ~500MB disk space (plus model storage)

More DevOps & CI/CD skills

Project scaffolding, deployment configuration, and CI/CD setup for Google ADK agents.

6K 472.6K
View

Set up tracing, logging, and monitoring for deployed ADK agents across Cloud Trace, BigQuery, and third-party platforms.

6K 472.6K
View

Enterprise Azure infrastructure architect generating Bicep or Terraform from workload descriptions.

1.5K 454.3K
View
azure-kubernetes logo
DevOps & CI/CD

azure-kubernetes

Plan and configure production-ready Azure Kubernetes Service clusters with Day-0 and Day-1 best practices.

1.5K 447.1K
View

Raw mechanical interfaces fusing Swiss typographic print with military terminal aesthetics. Rigid grids, extreme type scale contrast, utilitarian color, analog degradation effects. For data-heavy dashboards, portfolios, or editorial sites that need to feel like declassified blueprints.

92.7K 354.9K
View
just-scrape logo
DevOps & CI/CD

just-scrape

Web search, scraping, extraction, crawling, and monitoring via ScrapeGraph AI CLI.

69 245K
View

Skill for working with Firebase Hosting (Classic). Use this when you want to deploy static web apps, Single Page Apps (SPAs), or simple microservices. Do NOT use for Firebase App Hosting.

462 159.7K
View

Deploy and manage web apps with Firebase App Hosting. Use this skill when deploying Next.js/Angular apps with backends.

462 159.1K
View
deploy-to-vercel logo
DevOps & CI/CD

deploy-to-vercel

Deploy applications and websites to Vercel. Use when the user requests deployment actions like "deploy my app", "deploy and give me the link", "push this live", or "create a preview deployment".

31.9K 146.6K
View
programmatic-seo logo
DevOps & CI/CD

programmatic-seo

Build SEO-optimized pages at scale using templates, data, and proven playbook patterns.

53.3K 140.1K
View

Design and build isolated, reusable Convex backend components with clear boundaries and app-facing wrappers.

63 123.5K
View

Deploy and manage projects on Vercel using token-based authentication. Use when working with Vercel CLI using access tokens rather than interactive login — e.g. "deploy to vercel", "set up vercel", "add environment variables to vercel".

31.9K 116.1K
View

Coding & apps AI tools

Opus Clip logo
Coding & apps

Opus Clip

Opus.ai: Revolutionize Your Web Experience

Free
View
I
Coding & apps

Imagica

Build a no-code AI app in minutes.

Freemium
View
E
Coding & apps

Emergent.sh

An IDE for code migration from legacy to modern frameworks through coding agents.

Freemium
View
Wonder Dynamics logo
Coding & apps

Wonder Dynamics

Automate CGI animation in live-action scenes

Paid
View
M
Coding & apps

Mixo

Launch a website in seconds with AI.

Paid
View
AI Code Convert logo
Coding & apps

AI Code Convert

Streamline Your Coding Experience with AI Code Helper

Free
View