Downloads · 30 days
16
18% of all-time downloads
Rcarvalo/vibevoice
vibevoice is a text-to-speech model from Rcarvalo. Use it when you need text read aloud. It is set up for vibevoice. The card lists the license as mit.
Fine-tuned version of microsoft/VibeVoice-Realtime-0.5B on the French SIWIS dataset for improved French TTS.
Downloads · 30 days
16
18% of all-time downloads
All-time downloads
91
Public
Parameters
434M
1.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.7 GB · 100%
From the Hugging Face model README
Fine-tuned version of microsoft/VibeVoice-Realtime-0.5B on the French SIWIS dataset for improved French TTS.
| Parameter | Value |
|---|---|
| Epochs | 10 |
| Batch size | 4 |
| Gradient accumulation | 4 |
| Effective batch size | 16 |
| Learning rate | 5e-5 |
| Weight decay | 0.01 |
| Warmup steps | 500 |
| Precision | bf16 |
| Metric | Value |
|---|---|
| WER (mean) | 35.0% |
| WER (median) | 22.9% |
| RTF (mean) | 0.416 |
import torch
import soundfile as sf
from vibevoice.modular.modeling_vibevoice_streaming_inference import (
VibeVoiceStreamingForConditionalGenerationInference,
)
model = VibeVoiceStreamingForConditionalGenerationInference.from_pretrained(
"Rcarvalo/vibevoice",
torch_dtype=torch.bfloat16,
).to("cuda")
# Generate French speech
audio = model.generate(text="Bonjour, comment allez-vous aujourd'hui?")
sf.write("output.wav", audio.cpu().numpy(), 24000)
MIT (same as base model)