Downloads · 30 days
0
rootxhacker/Arthemis-TTS-22M
Arthemis-TTS-22M is a machine learning model from rootxhacker. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A lightweight transformer-based text-to-speech model that generates high-quality speech with a natural American women's voice. Despite having only 21-22 million parameters, this model delivers surprisingly clear synth…
Downloads · 30 days
0
Access
Public
Updated Sep 9, 2025
Repo size
1.1 GB
Likes
0
Public
Click a slice to open those files.
.pt1.1 GB · 100%
From the Hugging Face model README
A lightweight transformer-based text-to-speech model that generates high-quality speech with a natural American women's voice. Despite having only 21-22 million parameters, this model delivers surprisingly clear synthesis
This is a compact implementation of the Transformer TTS architecture, specifically trained to produce speech that sounds like a American woman speaker. The model strikes an excellent balance between medium quality and efficiency, making it perfect for applications where you need good speech synthesis without the computational overhead of larger models.
Key Features:
pip install arthemis-tts
import arthemis_tts
# Generate speech with American women's voice
audio = arthemis_tts.text_to_speech(
"Hello,world!",
model_path="arthemis_final.pt",
output_path="test_speech.wav"
)
This model uses a simplified Transformer TTS architecture with:
The architecture is intentionally kept lightweight while maintaining the core transformer mechanisms that make modern TTS systems effective.
This model is particularly well-suited for:
For detailed usage examples and advanced configuration options, please refer to the Arthemis TTS library documentation
This model is released under the MIT License. Feel free to use it in both commercial and non-commercial projects.
Note: This model represents a good starting point for English TTS applications. While it may not match the quality of much larger models (100M+ parameters), it offers an excellent trade-off between size, speed, and quality for most practical applications.
Happy synthesizing! 🎤