Downloads ยท 30 days
0
Expodecaprio/OmniVoice-Studio
OmniVoice-Studio is a machine learning model from Expodecaprio. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads ยท 30 days
0
Access
Public
Updated Jul 19, 2026
Repo size
โ
Likes
1
Public
Click a slice to open those files.
.py469 KB ยท 95%
From the Hugging Face model README
An enhanced, streamlined, and modular version of OmniVoice, featuring persistent .pt voice prompt serialization, clean English-only UI, custom offline caching, and optimized generation controls.
[!CAUTION]
๐ STRICT PROHIBITION ON UNAUTHORIZED VOICE CLONING
Do NOT use this software, model weights, or generated audio for any purpose involving cloning the voice of any individual without their explicit, verifiable, prior written consent.
By downloading, cloning, or operating this project, you strictly agree to the following terms:
- No Unauthorized Cloning: You will not impersonate any person (living or deceased), public figure, celebrity, or private individual without valid legal consent.
- No Deceptive Content: You will not generate defamatory, misleading, fraudulent, non-consensual sexual, or politically manipulative speech (Deepfakes).
- Non-Commercial Restriction: The pre-trained neural network weights (
k2-fsa/OmniVoice) are licensed strictly under Creative Commons Attribution-NonCommercial 4.0 International (CC-BY-NC 4.0). Commercial deployment, monetization, or paid API integration of the model weights is forbidden without explicit authorization from the original copyright holders.The creators and maintainers of this repository assume no liability for misuse, illegal activities, or copyright/personality rights violations arising from the operation of this software.
This project is built upon the foundational work of OmniVoice by the Xiaomi AI Lab Next-gen Kaldi Team. We give our utmost gratitude and full credit to the original authors and maintainers:
zhuhan), Next-gen Kaldi team (k2-fsa), and contributors.All core neural network architectures, flow-matching mechanisms, and multi-lingual acoustic modeling belong to their pioneering research.
This repository introduces several critical engineering upgrades to make local deployment cleaner, faster, and more reusable:
VoiceClonePrompt.save & .load)Instead of re-extracting speaker embeddings and processing reference audio on every single synthesis request, this enhanced version allows:
.pt): Extract the acoustic prompt (VoiceClonePrompt) once from any reference audio and save it to a lightweight .pt file..pt prompt file in the UI to skip reference audio processing entirely.saved_prompts/.0.9 for more natural, deliberate pacing and clearer articulation across complex sentences.HF_HOME, HF_HUB_CACHE) and PyTorch cache (TORCH_HOME) to a local models_cache/ directory when running locally on Windows/Linux.C:\ drive exhaustion when downloading the ~3.36 GB model weights.HF_HUB_OFFLINE=1 once weights are downloaded locally to eliminate startup latency and network timeouts.git clone https://github.com/Fazzbro/OmniVoice-Studio.git
cd OmniVoice-Studio
python -m venv .venv
# Windows PowerShell:
.\.venv\Scripts\Activate.ps1
# Linux / macOS:
source .venv/bin/activate
pip install --upgrade pip
pip install -r requirements.txt
Simply launch the entry point script:
python "Omni Voice.py"
k2-fsa/OmniVoice checkpoint (~3.36 GB) and store it safely in the models_cache/ folder.http://127.0.0.1:7860
OmniVoice-Studio/
โโโ app.py # Hugging Face Spaces entry point
โโโ Omni Voice.py # Local runner with D:\ drive offline cache config & Gradio server
โโโ requirements.txt # Clean, exact Python dependencies
โโโ LICENSE # Apache 2.0 License with Xiaomi AI Lab attribution
โโโ README.md # Project documentation, Spaces YAML, and legal disclaimer
โโโ omnivoice/ # Enhanced bundled package (self-contained)
โโโ __init__.py # Exports OmniVoice, VoiceClonePrompt, generation configs
โโโ models/
โ โโโ omnivoice.py # Core model + VoiceClonePrompt serialization (.save/.load)
โโโ cli/
โ โโโ demo.py # Streamlined English-only Gradio UI (0.9 speed default)
โโโ ... # Core utility and generation modules
k2-fsa/OmniVoice): Distributed under CC-BY-NC 4.0. Non-commercial use only.