Downloads · 30 days
0
oneplanettechnologies/my-big-model
my-big-model is a machine learning model from oneplanettechnologies. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A complete, production-ready Text-to-Speech (TTS) system that crafts human-like voices from text. This system uses Tacotron2 for mel-spectrogram generation and HiFiGAN for high-quality neural vocoding.
Downloads · 30 days
0
Access
Public
Updated Nov 29, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md24.4 KB · 74%
From the Hugging Face model README
A complete, production-ready Text-to-Speech (TTS) system that crafts human-like voices from text. This system uses Tacotron2 for mel-spectrogram generation and HiFiGAN for high-quality neural vocoding.
Voicecraft is a comprehensive TTS system developed by students of Dhemaji Engineering College, Computer Science and Engineering Department, under the guidance of Barga Deori Sir.
Team Members:
voicecraft/
├── datasets/ # Dataset handling
│ ├── download_ljspeech.py
│ └── prepare_dataset.py
├── preprocessing/ # Text and audio preprocessing
│ ├── text_normalizer.py
│ └── audio_processor.py
├── g2p/ # Grapheme-to-Phoneme conversion
│ └── g2p_converter.py
├── prosody/ # Prosody prediction
│ └── prosody_predictor.py
├── tacotron2/ # Tacotron2 acoustic model
│ ├── model.py
│ └── train.py
├── vocoder/ # Vocoders (HiFiGAN, WaveGlow)
│ ├── hifigan.py
│ └── train.py
├── utils/ # Utility functions
│ ├── audio_utils.py
│ ├── collate_fn.py
│ └── evaluation.py
├── deployment/ # Deployment files
│ ├── app.py # Gradio GUI
│ ├── requirements.txt
│ └── colab_deploy.ipynb
├── future/ # Future expansion modules
│ ├── assamese_tts.py
│ ├── emotional_tts.py
│ ├── whisper_transcription.py
│ └── mobile_inference.py
└── README.md
git clone <repository-url>
cd voicecraft
pip install -r deployment/requirements.txt
python voicecraft/datasets/download_ljspeech.py
python voicecraft/datasets/prepare_dataset.py data/LJSpeech-1.1
python voicecraft/deployment/app.py
Open your browser to http://localhost:7860
Use the interface:
# Download LJSpeech
python voicecraft/datasets/download_ljspeech.py
# Prepare dataset with train/val/test splits
python voicecraft/datasets/prepare_dataset.py data/LJSpeech-1.1
The dataset will be split into:
For experimental purposes, a small custom dataset (100-500 samples) can be used for GlowTTS experimentation.
Text Preprocessing
G2P Conversion
Prosody Prediction
Tacotron2 Acoustic Model
HiFiGAN Vocoder
Post-processing
python voicecraft/tacotron2/train.py \
--train_data data/processed/train \
--val_data data/processed/val \
--checkpoint_dir checkpoints/tacotron2 \
--log_dir logs/tacotron2 \
--num_epochs 100 \
--batch_size 32 \
--lr 1e-3 \
--use_amp \
--save_interval 1000
Training Parameters:
Loss Functions:
python voicecraft/vocoder/train.py \
--train_data data/processed/train \
--checkpoint_dir checkpoints/hifigan \
--log_dir logs/hifigan \
--num_epochs 100 \
--batch_size 16 \
--lr_g 2e-4 \
--lr_d 2e-4 \
--use_amp \
--save_interval 1000
Training Parameters:
MCD (Mel Cepstral Distortion)
Spectral Distortion
MOS Helper
from voicecraft.utils.evaluation import compute_mcd, compute_spectral_distortion
mcd = compute_mcd(mel_pred, mel_target)
sd = compute_spectral_distortion(mel_pred, mel_target)
Create a new Space on HuggingFace
Upload files:
voicecraft/deployment/app.pyvoicecraft/deployment/requirements.txtvoicecraft/ source codeConfigure Space:
The app will automatically:
Upload the notebook:
voicecraft/deployment/colab_deploy.ipynbRun cells:
Create public tunnel:
from pyngrok import ngrok
public_url = ngrok.connect(7860)
print(f"Public URL: {public_url}")
Share the link for public access
# Install dependencies
pip install -r voicecraft/deployment/requirements.txt
# Run the app
python voicecraft/deployment/app.py
The app will be available at http://localhost:7860
The Gradio interface includes:
Click the "📋 Paste Intro Paragraph" button to automatically fill the text box with:
Hello everyone, hope you are well. We are students of Dhemaji Engineering College, Computer Science and Engineering Department, and our names are: Abinash Dutta, Arindom Bordoloi, Bakhtiar Ahmed, Chinmoy Mahanta, Mumon Saikia, Preety Rani Boruah. We are developing the project "Voicecraft: The Art of Text-to-Speech – Crafting Human-Like Voices from Text" under the guidance of Barga Deori Sir. Thank you.
Assamese TTS (future/assamese_tts.py)
Emotional TTS (future/emotional_tts.py)
Whisper Transcription (future/whisper_transcription.py)
Mobile Inference (future/mobile_inference.py)
voicecraft/datasets/voicecraft/preprocessing/voicecraft/g2p/voicecraft/tacotron2/ and voicecraft/vocoder/voicecraft/deployment/from voicecraft.preprocessing.text_normalizer import normalize_text
from voicecraft.g2p.g2p_converter import text_to_phonemes
from voicecraft.tacotron2.model import Tacotron2
from voicecraft.vocoder.hifigan import HiFiGAN
# Normalize text
text = "Hello, it's 9:30 a.m. on Jan. 15, 2023."
normalized = normalize_text(text)
# Convert to phonemes
phonemes = text_to_phonemes(normalized)
# Load models
tacotron2 = Tacotron2()
hifigan = HiFiGAN()
# Generate speech
text_seq = text_to_sequence(normalized)
mel = tacotron2.infer(text_seq)
audio = hifigan.infer(mel)
preprocessing/text_normalizer.pyg2p/g2p_converter.pytacotron2/ or vocoder/deployment/app.py# Test individual modules
python voicecraft/preprocessing/text_normalizer.py
python voicecraft/g2p/g2p_converter.py
[Add your license here]
For questions, issues, or contributions, please contact the development team.
Voicecraft: The Art of Text-to-Speech - Crafting Human-Like Voices from Text