Downloads ยท 30 days
20
13% of all-time downloads
useclaude/quantum-sync-xtts-v2
quantum-sync-xtts-v2 is a text-to-speech model from useclaude. Use it when you need text read aloud. It is set up for transformers. The card lists the license as mpl-2.0.
This is a mirror/backup of the Coqui XTTS-v2 model for use with the Quantum Sync project.
Downloads ยท 30 days
20
13% of all-time downloads
All-time downloads
152
Public
Repo size
1.9 GB
Likes
4
Public
Click a slice to open those files.
.pth1.9 GB ยท 100%
From the Hugging Face model README
This is a mirror/backup of the Coqui XTTS-v2 model for use with the Quantum Sync project.
This mirror serves as:
Original Model: coqui/XTTS-v2
Architecture: XTTS-v2 (Zero-shot multi-lingual TTS)
Model Size: ~1.87 GB
Supported Languages: 13 languages
git clone https://github.com/Useforclaude/quantum-sync-v5.git
cd quantum-sync-v5/quantum-sync-v11-production
# Configure to use this mirror
# Edit tts_engines/xtts.py, change model_name to:
# model_name = "useclaude/quantum-sync-xtts-v2"
python main_v11.py input/file.srt \
--voice MyVoice \
--voice-sample /path/to/voice.wav \
--tts-engine xtts-v2 \
--tts-language en
from TTS.api import TTS
# Use this mirror
tts = TTS(model_name="useclaude/quantum-sync-xtts-v2")
# Generate speech
tts.tts_to_file(
text="Hello, this is a test.",
speaker_wav="reference_voice.wav",
language="en",
file_path="output.wav"
)
from TTS.api import TTS
# Initialize
tts = TTS(model_name="useclaude/quantum-sync-xtts-v2")
# Clone voice from reference audio (6-30 seconds)
tts.tts_to_file(
text="The quick brown fox jumps over the lazy dog.",
speaker_wav="my_voice_sample.wav", # Your voice reference
language="en",
file_path="output_cloned.wav"
)
From Quantum Sync Production Tests (2025-10-13):
| Metric | Value |
|---|---|
| Synthesis Speed | ~3.7 segments/minute |
| Processing Time | 17 min for 277 segments (23 min audio) |
| Duration Accuracy | ~87% audio, ~13% silence gaps |
| Timeline Drift | -1.7% (excellent) |
| Voice Quality | 8/10 |
| Cloning Accuracy | Excellent |
| VRAM Usage | 6-8 GB |
Comparison:
# Speed control (0.5 - 2.0)
tts.tts_to_file(
text="Hello world",
speaker_wav="voice.wav",
language="en",
speed=0.8, # Slower speech
file_path="output.wav"
)
# Temperature control (0.1 - 1.0)
tts.tts_to_file(
text="Hello world",
speaker_wav="voice.wav",
language="en",
temperature=0.75, # More expressive
file_path="output.wav"
)
quantum-sync-xtts-v2/
โโโ model.pth (1.87 GB - Neural network weights)
โโโ config.json (Model configuration)
โโโ vocab.json (Vocabulary for tokenization)
โโโ speakers_xtts.pth (Speaker embeddings)
โโโ dvae.pth (DVAE component)
โโโ mel_stats.pth (Mel-spectrogram statistics)
โโโ LICENSE (MPL 2.0)
โโโ README.md (This file)
Mozilla Public License 2.0 (MPL 2.0)
This model is licensed under the Mozilla Public License 2.0. You can:
Requirements:
Full License: LICENSE
Original Work:
This Mirror:
All credit goes to the original Coqui TTS team. This is simply a mirror for backup and convenience.
Quantum Sync Documentation:
Original Documentation:
This is an unofficial mirror maintained for backup purposes. For the latest version and official support, please refer to the original model and Coqui TTS repository.
XTTS-v2 is a state-of-the-art zero-shot multi-lingual text-to-speech model that can clone voices from short audio samples (6-30 seconds).
Key Features:
Primary Use Cases:
Out-of-Scope Use:
XTTS-v2 was trained on diverse multi-lingual speech data. For details, see the original model card.
See Performance section above for detailed benchmarks from Quantum Sync project.
Voice Cloning Ethics:
Model Type: Autoregressive Transformer-based TTS
Framework: PyTorch
Input: Text + Reference Audio (6-30 sec WAV)
Output: 24kHz WAV audio
Inference Time: ~3-5 seconds per segment (GPU)
Hardware Requirements:
Software Requirements:
For this mirror:
For original model:
Last Updated: 2025-10-13
Mirror Version: 1.0
Model Version: XTTS-v2 (Latest as of upload date)