Downloads · 30 days
0
marcosremar2/ai-video-setup
ai-video-setup is a machine learning model from marcosremar2. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Complete setup for real-time AI video avatars with voice cloning and lip sync.
Downloads · 30 days
0
Access
Public
Updated Dec 11, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py18.8 KB · 63%
From the Hugging Face model README
Complete setup for real-time AI video avatars with voice cloning and lip sync.
| Component | RTF | Speed |
|---|---|---|
| StyleTTS2 (steps=5) | 0.04 | 22x real-time |
| MuseTalk V1.5 | 0.25-0.67 | 1.5-4x real-time |
# Clone and install
git clone https://github.com/yourusername/ai-video-setup.git
cd ai-video-setup
chmod +x install.sh
./install.sh
# Generate audio with cloned voice
python scripts/generate_audio.py --text "Hello world" --voice voice_ref.wav -o output.wav
# Run lip sync
python scripts/run_lipsync.py --video avatar.mp4 --audio output.wav -o ./output
# Or run the full pipeline
python scripts/full_pipeline.py --youtube-url "https://..." --text "Your text" -o ./output
These versions are required and tested to work together:
accelerate==0.25.0
diffusers==0.21.0
huggingface-hub==0.25.0
Warning: Newer versions cause
cannot import clear_device_cacheerror.
| Script | Description |
|---|---|
install.sh | Complete installation (PyTorch, MuseTalk, StyleTTS2) |
scripts/generate_audio.py | Generate audio with voice cloning |
scripts/run_lipsync.py | Run lip sync on video |
scripts/extract_voice_ref.py | Extract voice reference from YouTube/video |
scripts/full_pipeline.py | Complete pipeline (YouTube -> lip sync video) |
scripts/realtime_avatar.py | Real-time avatar with pre-loaded models |
For best lip sync results:
# Convert 4K video to optimal format
ffmpeg -i input_4k.mp4 -vf "scale=640:-2,fps=24" -c:a copy avatar.mp4
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ Input Text │────▶│ StyleTTS2 │────▶│ Audio WAV │
└─────────────────┘ │ (Voice Clone) │ └────────┬────────┘
└─────────────────┘ │
┌─────────────────┐ │
│ Avatar Video │──────────────────────────────────────┤
└─────────────────┘ │
┌─────────────────┐ │
│ MuseTalk V1.5 │◀─────────────┘
│ (Lip Sync) │
└────────┬────────┘
│
┌────────▼────────┐
│ Output Video │
│ (with audio) │
└─────────────────┘
pip install accelerate==0.25.0 diffusers==0.21.0 huggingface-hub==0.25.0
The scripts include a fix for weights_only parameter. If you see pickle errors, ensure the patch is applied:
import torch
original_load = torch.load
def patched_load(*args, **kwargs):
kwargs['weights_only'] = False
return original_load(*args, **kwargs)
torch.load = patched_load
python -c "import nltk; nltk.download('punkt'); nltk.download('punkt_tab')"
This project integrates: