Downloads · 30 days
16
12% of all-time downloads
developerabu/vits-tts-mnn
vits-tts-mnn is a text-to-speech model from developerabu. Use it when you need text read aloud. It is set up for transformers. The card lists the license as cc-by-4.0.
This repository contains a VITS-based Text-to-Speech (TTS) model fine-tuned for Indian languages. The model supports multiple Indian languages and a wide range of speaking styles and emotions, making it suitable for d…
Downloads · 30 days
16
12% of all-time downloads
All-time downloads
139
Public
Repo size
220 MB
Likes
2
Public
Click a slice to open those files.
.onnx187 MB · 85%
From the Hugging Face model README
This repository contains a VITS-based Text-to-Speech (TTS) model fine-tuned for Indian languages. The model supports multiple Indian languages and a wide range of speaking styles and emotions, making it suitable for diverse use cases such as conversational AI, audiobooks, and more.
The model ai4bharat/vits_rasa_13 is based on the VITS architecture and supports the following features:
pip install transformers torch
Here's a quick example to get started:
import soundfile as sf
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained("ai4bharat/vits_rasa_13", trust_remote_code=True).to("cuda")
tokenizer = AutoTokenizer.from_pretrained("ai4bharat/vits_rasa_13", trust_remote_code=True)
text = "ਕੀ ਮੈਂ ਇਸ ਹਫਤੇ ਦੇ ਅੰਤ ਵਿੱਚ ਰੁੱਝਿਆ ਹੋਇਆ ਹਾਂ?" # Example text in Punjabi
speaker_id = 16 # PAN_M
style_id = 0 # ALEXA
inputs = tokenizer(text=text, return_tensors="pt").to("cuda")
outputs = model(inputs['input_ids'], speaker_id=speaker_id, emotion_id=style_id)
sf.write("audio.wav", outputs.waveform.squeeze(), model.config.sampling_rate)
print(outputs.waveform.shape)
AssameseBengaliBodoDogriKannadaMaithiliMalayalamMarathiNepaliPunjabiSanskritTamilTeluguIf you use this model in your research, please cite:
@article{ai4bharat_vits_rasa_13,
title={VITS TTS for Indian Languages},
author={Ashwin Sankar},
year={2024},
publisher={Hugging Face}
}