Downloads ยท 30 days
27
26% of all-time downloads
FenomAI/OmniVoice
OmniVoice is a text-to-speech model from FenomAI. Use it when you need text read aloud. It is set up for omnivoice. The card lists the license as apache-2.0.
<p align="center" <img width="200" height="200" alt="OmniVoice" src="https://zhu-han.github.io/omnivoice/pics/omnivoice.jpg" / </p
Downloads ยท 30 days
27
26% of all-time downloads
All-time downloads
103
Public
Parameters
613M
3.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors3.3 GB ยท 100%
How the weights are stored.
F32613M ยท 100%
From the Hugging Face model README
OmniVoice is a massively multilingual zero-shot text-to-speech (TTS) model supporting over 600 languages. Built on a novel diffusion language model-style architecture, it delivers high-quality speech with superior inference speed, supporting voice cloning and voice design.
[laughter]) and pronunciation correction via pinyin or phonemes.To get started, install the omnivoice library:
We recommend using a fresh virtual environment (e.g.,
conda,venv, etc.) to avoid conflicts.
Step 1: Install PyTorch
<details> <summary>NVIDIA GPU</summary># Install pytorch with your CUDA version, e.g.
pip install torch==2.8.0+cu128 torchaudio==2.8.0+cu128 --extra-index-url https://download.pytorch.org/whl/cu128
</details> <details> <summary>Apple Silicon</summary>See PyTorch official site for other versions installation.
pip install torch==2.8.0 torchaudio==2.8.0
</details>
Step 2: Install OmniVoice
pip install omnivoice
You can use OmniVoice for zero-shot voice cloning as follows:
from omnivoice import OmniVoice
import soundfile as sf
import torch
# Load the model
model = OmniVoice.from_pretrained(
"k2-fsa/OmniVoice",
device_map="cuda:0",
dtype=torch.float16
)
# Generate audio
audio = model.generate(
text="Hello, this is a test of zero-shot voice cloning.",
ref_audio="ref.wav",
ref_text="Transcription of the reference audio.",
) # audio is a list of `np.ndarray` with shape (T,) at 24 kHz.
sf.write("out.wav", audio[0], 24000)
For more generation modes (e.g., voice design), functions (e.g., non-verbal symbols, pronunciation correction) and comprehensive usage instructions, see our GitHub Repository.
You can directly discuss on GitHub Issues.
You can also scan the QR code to join our wechat group or follow our wechat official account.
| Wechat Group | Wechat Official Account |
|---|---|
![]() | ![]() |
@article{zhu2026omnivoice,
title={OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models},
author={Zhu, Han and Ye, Lingxuan and Kang, Wei and Yao, Zengwei and Guo, Liyong and Kuang, Fangjun and Han, Zhifeng and Zhuang, Weiji and Lin, Long and Povey, Daniel},
journal={arXiv preprint arXiv:2604.00688},
year={2026}
}