Downloads ยท 30 days
0
ViuAI/ViuAI_TTS_200M
ViuAI_TTS_200M is a text-to-speech model from ViuAI. Use it when you need text read aloud. The card lists the license as apache-2.0.
ViuAITTS200M is a state-of-the-art (SOTA) Text-to-Speech deep learning architecture designed for human-grade speech synthesis and zero-shot voice cloning.
Downloads ยท 30 days
0
Access
Public
Updated Sep 22, 2026
Repo size
32.3 GB
Likes
0
Public
Click a slice to open those files.
.pt33 GB ยท 100%
From the Hugging Face model README
ViuAI_TTS_200M is a state-of-the-art (SOTA) Text-to-Speech deep learning architecture designed for human-grade speech synthesis and zero-shot voice cloning.
Inspired by modern breakthrough architectures (Flow Matching, DiT, ElevenLabs Adam, and Grok Voice), ViuAI_TTS_200M eliminates robotic artifacts through:
--ref_audio sample.wav) to replicate any target voice's timbre, warmth, and cadence.| Component | Architecture Specs | Trainable Parameters |
|---|---|---|
| Text & Prosody Encoder | Pre-LN Transformer + RoPE (Multilingual) | 19.43 Million |
| Human Prosody Engine | Dynamic $F_0$ Pitch + Energy + Duration ConvNeXt | 8.42 Million |
| DiT Core Backbone | 9-layer Diffusion Transformer + AdaLN-Zero + CFG | 148.71 Million |
| Neural Audio Vocoder | 24,000 Hz Multi-Receptive Periodic Generator | 13.93 Million |
| Total Model Parameters | 190.49 Million (~190.5M) |
Clone any human voice by supplying a 3โ5 second .wav audio prompt:
python inference.py \
--text "Hello! This is ViuAI_TTS_200M speaking in a human-like voice." \
--ref_audio "path/to/reference_sample.wav" \
--cfg_scale 2.0 \
--steps 15 \
--output "cloned_output.wav"
python inference.py \
--text "Namaste! Yeh ViuAI TTS 200M model ka audio output hai." \
--cfg_scale 2.0 \
--steps 15 \
--output "speech.wav"
Train with Mixed Precision (AMP), AdamW, and Cosine Annealing:
python train.py
Checkpoints are automatically saved to checkpoints/ and can be loaded directly for inference.
ViuAI_TTS_200M/
โโโ config/
โ โโโ model_config.py # 190.5M architectural hyperparameters
โโโ models/
โ โโโ text_encoder.py # Multilingual Transformer with RoPE
โ โโโ dit_backbone.py # Diffusion Transformer with Voice Conditioning
โ โโโ duration_predictor.py # Human Prosody Engine (F0, Energy, Duration)
โ โโโ flow_matching.py # Optimal Transport CFM & Euler ODE Solver
โ โโโ vocoder.py # 24kHz Neural Audio Vocoder
โ โโโ viuai_tts.py # Unified ViuAITTS200M Model
โโโ dataset/
โ โโโ audio_processing.py # Mel-spectrogram extraction
โ โโโ tokenizer.py # Hindi & English multilingual tokenizer
โโโ checkpoints/
โ โโโ viuai_tts_200m_init.pt # Pretrained / initialized model weights
โโโ train.py # Mixed Precision (AMP) training script
โโโ inference.py # Text-to-speech audio generator
โโโ verify_model.py # Architecture verification script
โโโ ROADMAP.md # Master engineering roadmap
Apache 2.0