Downloads · 30 days
20
11% of all-time downloads
AEmotionStudio/omnivoice-models
omnivoice-models is a text-to-speech model from AEmotionStudio. Use it when you need text read aloud. The card lists the license as apache-2.0.
Multi-Lingual TTS & Voice Cloning — 600+ Languages
Downloads · 30 days
20
11% of all-time downloads
All-time downloads
178
Public
Parameters
613M
3.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.3 GB · 100%
How the weights are stored.
F32613M · 100%
From the Hugging Face model README
Multi-Lingual TTS & Voice Cloning — 600+ Languages
Original Model by k2-fsa (Next-gen Kaldi) · Apache 2.0
This is a mirror of the OmniVoice model weights for use with Mæstræa AI Workstation. All credits go to the original authors.
| Path | Description | Size |
|---|---|---|
model.safetensors | Main OmniVoice model | ~3 GB |
audio_tokenizer/model.safetensors | Audio tokenizer | ~260 MB |
tokenizer.json | Text tokenizer | ~17 MB |
config.json | Model configuration | < 1 KB |
OmniVoice is a multi-lingual TTS and voice cloning model supporting 600+ languages with near real-time inference (RTF ~0.025). It supports three modes:
These models are automatically downloaded by the Mæstræa AI Workstation backend. They can also be loaded manually:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("AEmotionStudio/omnivoice-models")
tokenizer = AutoTokenizer.from_pretrained("AEmotionStudio/omnivoice-models")
Apache 2.0 — same as the original OmniVoice release.