Downloads · 30 days
975
35% of all-time downloads
xmanii/Ava-82M
Ava-82M is a text-to-speech model from xmanii. Use it when you need text read aloud. It is set up for kokoro. The card lists the license as apache-2.0.
Ava is a lightweight, open-weight Persian text-to-speech model released by Nimruz. It is a Persian adaptation of Kokoro-82M, fine-tuned on an approximately 20-hour curated subset of Mana-TTS.
Downloads · 30 days
975
35% of all-time downloads
All-time downloads
2.8K
Public
Repo size
657 MB
Likes
18
Public
Click a slice to open those files.
.pth327 MB · 99%
From the Hugging Face model README
Ava is a lightweight, open-weight Persian text-to-speech model released by Nimruz. It is a Persian adaptation of Kokoro-82M, fine-tuned on an approximately 20-hour curated subset of Mana-TTS.
Ava produces 24 kHz, single-speaker Persian speech. v0.2 continues training from the listener-selected v0.1 model on approximately 20 total hours of Mana-TTS. The release includes a Persian text frontend, number/date normalization, contextual grapheme-to-phoneme conversion, pronunciation overrides, automatic long-text splitting, and conservative cleanup of leading and trailing synthesis artifacts.
Status: v0.2 research release. Informal native-speaker listening found improved tone, pace, pronunciation, and stability over v0.1. Occasional robotic pitch and synthesis artifacts remain; no formal MOS or Persian intelligibility benchmark has been completed yet.
Python 3.11–3.13 and Git are required. Install the small Ava package directly from this repository:
python -m pip install \
"https://huggingface.co/xmanii/Ava-82M/resolve/main/ava_tts-0.2.0-py3-none-any.whl"
Then generate speech:
from ava_tts import Ava
tts = Ava()
tts.save("سلام! من آوا هستم.", "ava.wav")
The first run downloads the 82M-parameter acoustic model and the pinned Persian G2P frontend. Later runs use the local cache.
ava "قیمت این محصول سه میلیون و چهارصد هزار تومان است." -o price.wav
Adjust speaking rate:
ava "امروز بیست و ششم تیر است." --speed 0.95 -o slower.wav
Boundary cleanup is enabled by default because listening tests found that it removes the short robotic noise that can occur before or after speech. To inspect the unprocessed model output:
ava "سلام، این خروجی خام مدل است." --raw-boundaries -o raw.wav
Loading the model once is faster:
from ava_tts import Ava
tts = Ava.from_pretrained("xmanii/Ava-82M")
tts.save("این جملهٔ اول است.", "first.wav")
tts.save("و این جملهٔ دوم است.", "second.wav", speed=1.05)
On the tested MacBook Air M4, CPU inference generated approximately 4.5 seconds of speech in roughly 0.9 seconds after loading. CPU is therefore selected automatically on Apple Silicon.
| File | Purpose |
|---|---|
model.pth | Selected 20-hour Stage-2 Epoch-1 acoustic checkpoint |
voice.pt | Persian single-speaker style/voice tensor |
config.json | Kokoro architecture and phoneme vocabulary |
pronunciations.json | Reviewed pronunciation correction layer |
ava_tts-0.2.0-py3-none-any.whl | Friendly Python API and ava command |
samples/ava-v02-*.wav | Boundary-cleaned v0.2 examples |
artifact_manifest.json | File sizes, hashes, and provenance |
The package pins the exact Kokoro runtime revision used during local and cloud
validation. The public kokoro==0.9.4 decoder is not weight-compatible with
this checkpoint.
Ava retains Kokoro's compact StyleTTS2/iSTFTNet-derived inference architecture:
ژ errors, digit-nine ambiguity, and
ezafe liaison before mapping phonemes into Kokoro's IPA vocabulary.The acoustic network has 81,763,410 inference parameters. The separate G2P frontend is not included in that parameter count.
hexgrad/Kokoro-82MMahtaFetrat/Mana-TTS0.4430660.5462.880The Stage-2 epoch completed 3,850 batch-4 updates on an L40S. All tensors in the saved checkpoint passed a direct finiteness audit. Joint adversarial training remains disabled because the earlier 10-hour experiment became numerically unstable and its listener-evaluated output was more robotic.
Evaluation for v0.2 is primarily structured human listening:
ژ, digit nine, and ordinal dates
without modifying the acoustic weights.Formal MOS, speaker-similarity, real-time-factor across platforms, and Persian phoneme-error-rate evaluations remain future work.
main / v0.2.0: 20-hour Stage-2 Epoch-1 model and frontend v3.v0.1 branch: preserved final v0.1 repository state.v0.1.0 and v0.1.1 tags: original 10-hour releases.Use Ava for ethical speech and accessibility applications. Do not use it for impersonation, identity theft, fraud, deception, harassment, or misleading claims that generated speech came from a real person. Clearly disclose synthetic speech where a listener could reasonably mistake it for authentic recorded speech.
The Mana-TTS maintainers specifically prohibit impersonation, identity theft, and fraudulent use in their ethical-use notice.
Ava code and weights are released under Apache License 2.0. See
LICENSE and NOTICE.md.
This release depends on:
If you use Ava, please cite this repository as well as Kokoro, StyleTTS2,
Mana-TTS, and Homo-GE2PE. Full upstream citations are included in
CITATIONS.md.