Downloads · 30 days
52
12% of all-time downloads
AEmotionStudio/fish-speech-s2-pro
fish-speech-s2-pro is a text-to-speech model from AEmotionStudio. Use it when you need text read aloud. The card lists the license as other.
Mirror of the Fish Speech S2 Pro model by Fish Audio.
Downloads · 30 days
52
12% of all-time downloads
All-time downloads
423
Public
Parameters
4.6B
11 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors9.1 GB · 83%
From the Hugging Face model README
Mirror of the Fish Speech S2 Pro model by Fish Audio.
Original model: fishaudio/fish-speech-1.5
| File | Size | Description |
|---|---|---|
model.safetensors | 9.12 GB | Main language model weights |
codec.pth | 1.87 GB | Audio codec (encoder/decoder) |
config.json | 1.86 KB | Model configuration |
tokenizer.json | 12.2 MB | Tokenizer data |
tokenizer_config.json | 861 KB | Tokenizer configuration |
special_tokens_map.json | 102 KB | Special tokens mapping |
chat_template.jinja | 4.12 KB | Chat template |
Fish Speech is a leading open-source text-to-speech (TTS) model that supports high-quality voice cloning and multilingual speech synthesis. The S2 Pro variant offers improved quality and zero-shot voice cloning capabilities.
This model is automatically downloaded and used by the ComfyUI-FFMPEGA extension for TTS and voice cloning features.
Fish Audio Research License — see LICENSE file.