Downloads · 30 days
17
30% of all-time downloads
AEmotionStudio/soulx-singer-models
soulx-singer-models is a text-to-speech model from AEmotionStudio. Use it when you need text read aloud. The card lists the license as apache-2.0.
Safetensors conversion of Soul-AILab/SoulX-Singer weights for use in the MAESTRO AI Workstation.
Downloads · 30 days
17
30% of all-time downloads
All-time downloads
57
Public
Repo size
5.9 GB
Likes
1
Public
Click a slice to open those files.
.safetensors5.9 GB · 100%
From the Hugging Face model README
Safetensors conversion of Soul-AILab/SoulX-Singer weights for use in the MAESTRO AI Workstation.
| Path | Size | Description |
|---|---|---|
| svs/model.safetensors | ~2.82 GB | Singing Voice Synthesis (lyrics+MIDI → singing) |
| svc/model.safetensors | ~2.79 GB | Singing Voice Conversion (audio-to-audio) |
| whisper-base/ | ~0.29 GB | Verbatim openai/whisper-base (Apache-2.0) — the SVC model's frozen semantic encoder. Its weights are not in the SVC checkpoint; bundling them here keeps SVC fully offline. |
| config.yaml | 579 B | Model architecture configuration |
| phone_set.json | ~30 KB | Phoneme mapping for SVS |
The SVS annotation pipeline (vocal separation → RMVPE F0 → lyric ASR → ROSVOT note transcription) pulls its component weights on demand from the companion repo AEmotionStudio/soulx-singer-preprocess (verbatim upstream files from the SoulX-Singer-Preprocess release).
svs/ and svc/ are torch.load(ckpt)["state_dict"] → safetensors
conversions of upstream model.pt / model-svc.pt; loaders consume them
with strict=True.whisper-base/ is byte-identical to upstream openai/whisper-base
(sha256-verified at upload; see
backend/scripts/mirror_soulx_singer_to_hf.py in MAESTRO).Apache 2.0