Downloads · 30 days
10
6% of all-time downloads
AsmaaAbdelkader/CodeSwitching_TTS_finetuned
CodeSwitching_TTS_finetuned is a text-to-speech model from AsmaaAbdelkader. Use it when you need text read aloud. The card lists the license as apache-2.0.
This is a fine-tuned XTTS model for Arabic Text-to-Speech. It has been trained to generate highly natural Arabic speech, taking advantage of a reference audio file for zero-shot voice cloning.
Downloads · 30 days
10
6% of all-time downloads
All-time downloads
173
Public
Repo size
5.6 GB
Likes
0
Public
Click a slice to open those files.
.pth5.6 GB · 100%
From the Hugging Face model README
This is a fine-tuned XTTS model for Arabic Text-to-Speech. It has been trained to generate highly natural Arabic speech, taking advantage of a reference audio file for zero-shot voice cloning.
You can easily use this model in Python. The following code downloads the model.pth, config.json, and vocab.json directly from this repository, initializes the model, and safely processes long Arabic text by splitting it into smaller sentence chunks.
This model was trained and tested using Python 3.10 (specifically 3.10.18). To run the inference code, you will need to install system-level audio libraries and specific Python dependencies.
1. Install System Audio Dependencies (Linux/Ubuntu): You need FFmpeg to handle audio processing natively. Run this in your terminal:
sudo apt update
sudo apt install ffmpeg libavcodec-dev libavutil-dev libavformat-dev
2. Install Python Dependencies: Install the core TTS library directly from the Coqui GitHub repository, along with the specific version of Transformers required for compatibility:
pip install git+[https://github.com/coqui-ai/TTS](https://github.com/coqui-ai/TTS)
pip install transformers==4.33.0
pip install torchcodec torch torchaudio huggingface_hub
import os
import torch
import torchaudio
import re
from TTS.tts.configs.xtts_config import XttsConfig
from TTS.tts.models.xtts import Xtts
from huggingface_hub import hf_hub_download
# --- Configuration ---
REPO_ID = "AsmaaAbdelkader/CodeSwitching_TTS_finetuned"
print("Step 1: Downloading model files from Hugging Face...")
config_path = hf_hub_download(repo_id=REPO_ID, filename="config.json")
vocab_path = hf_hub_download(repo_id=REPO_ID, filename="vocab.json")
checkpoint_path = hf_hub_download(repo_id=REPO_ID, filename="model.pth")
# Users can replace this with their own local file path if they want to clone a different voice
reference_audio = hf_hub_download(repo_id=REPO_ID, filename="reference_audio.wav")
print("Step 2: Loading the fine-tuned model...")
config = XttsConfig()
config.load_json(config_path)
model = Xtts.init_from_config(config)
model.load_checkpoint(
config,
checkpoint_path=checkpoint_path,
vocab_path=vocab_path,
use_deepspeed=False
)
model.cuda()
print("Step 3: Computing speaker latents...")
gpt_cond_latent, speaker_embedding = model.get_conditioning_latents(audio_path=[reference_audio])
# --- Text to Synthesize ---
full_text = "المستخدم بيرد بتخمين، ووكيل الذكاء الاصطناعي بيبعت إشارة تانية، وبعدين في مرحلة ما، المستخدم ممكن يكسب اللعبة لو خمن صح. كنا مهتمين جداً ببحث اتجاه التواصل، لأن فيه سيناريوهات بيبقى لازم فيها الإنسان هو اللي يقدم الإشارات لوكيل الذكاء الاصطناعي."
print("Step 4: Processing text and generating audio chunks...")
# Split text into sentences using common Arabic/English punctuation
sentences = re.split(r'(?<=[.؟!])\s+', full_text)
wav_chunks = []
print(f"Total sentences to process: {len(sentences)}")
for i, sentence in enumerate(sentences):
if len(sentence.strip()) <= 1: # Skip empty strings
continue
print(f" Generating chunk {i+1}...")
out = model.inference(
text=sentence.strip(),
language="ar",
gpt_cond_latent=gpt_cond_latent,
speaker_embedding=speaker_embedding,
temperature=0.7,
)
wav_chunks.append(torch.tensor(out["wav"]))
print("Step 5: Saving final audio...")
if wav_chunks:
final_wav = torch.cat(wav_chunks, dim=0)
# Sample rate for XTTS is typically 24000
output_filename = "output_full.wav"
torchaudio.save(output_filename, final_wav.unsqueeze(0), 24000)
print(f"Success! Audio saved locally as {output_filename}")