Downloads · 30 days
0
Mmoa/tts711
tts711 is a machine learning model from Mmoa. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This project provides a REST API for Microsoft Edge's text-to-speech service with word-by-word subtitle generation and highlighting capabilities.
Downloads · 30 days
0
Access
Public
Updated Nov 7, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.html9 KB · 52%
From the Hugging Face model README
This project provides a REST API for Microsoft Edge's text-to-speech service with word-by-word subtitle generation and highlighting capabilities.
pip install -r requirements.txt
Start the Flask server:
python app.py
The API will be available at http://localhost:7860
POST /tts
Convert text to speech with word-by-word subtitles.
Request Body:
{
"text": "Text to convert to speech",
"voice": "en-US-JennyNeural" // Optional, defaults to en-US-JennyNeural
}
Response:
{
"audio": "base64_encoded_audio_data",
"subtitles": "SRT formatted subtitles",
"format": "mp3"
}
GET /tts/audio/<text>
Get only the audio file for the provided text.
Query Parameters:
voice (optional): Voice to use for TTSGET /tts/subtitles/<text>
Get only the subtitles for the provided text.
Query Parameters:
voice (optional): Voice to use for TTSGET /voices
Get a list of all available voices.
const text = "Hello, this is a test.";
const response = await fetch('http://localhost:7860/tts', {
method: 'POST',
headers: {
'Content-Type': 'application/json'
},
body: JSON.stringify({text: text})
});
const data = await response.json();
const audio = new Audio('data:audio/mpeg;base64,' + data.audio);
audio.play();
To deploy this API to Hugging Face Spaces:
This project is licensed under the MIT License.