Downloads · 30 days
15
27% of all-time downloads
ehabnegm/s2pro-egy
s2pro-egy is a text-to-speech model from ehabnegm. Use it when you need text read aloud. The card lists the license as other.
Merged production model + full training/serving toolkit. Phase 1 complete (2026-07-18/19); Phase 2 planned — see below.
Downloads · 30 days
15
27% of all-time downloads
All-time downloads
56
Public
Parameters
4.6B
12.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors10 GB · 80%
How the weights are stored.
BF164.1B · 91%
From the Hugging Face model README
Merged production model + full training/serving toolkit. Phase 1 complete (2026-07-18/19); Phase 2 planned — see below.
⚠️ License: derivative of
fishaudio/s2-pro(Fish Audio Research License — non-commercial only). Production/commercial use requires a written license from Fish Audio ([email protected]). Keep this repo private.
| Path | Contents |
|---|---|
/ (root) | Merged model, HF fish_qwen3_omni layout: safetensors shards + index, config.json, tokenizer, chat_template.jinja, codec.pth — loads in fish-speech main AND sglang-omni |
checkpoints/ | Phase-1 LoRA checkpoints (fast-AR only, steps 100–1200; step_000001200.ckpt is the one merged) |
scripts/ | Full pipeline: data prep, relabeling, eval (synth + Soniox judge), serving (PyTorch + sglang shim), weight conversion, patches |
configs/ | Training configs (text2semantic_finetune_egy2.yaml = the recipe that worked) + LoRA configs |
eval/ | Mega-paragraph A/B wavs + Soniox WER reports (baseline vs checkpoints) |
logs/ | Training log + tensorboard events |
| Dataset (HF, private) | Clips | Hours | Transcript source |
|---|---|---|---|
ehabnegm/noselleel-egyptian-tts | 8,766 | 26.2 | Soniox stt-async-v5 (6,344 clips, transcripts_soniox/train.jsonl) > text_raw (pre-CATT Whisper) |
ehabnegm/eqkawkab-egyptian-tts | 2,062 | 8.0 | text_raw (pre-CATT Whisper) |
ehabnegm/moustafa-sadek-egyptian-tts | 2,883 | 8.0 | Deepgram nova-3 |
ehabnegm/mosaifside-egyptian-tts | 1,319 | 4.7 | Deepgram nova-3 |
Processing rules (see scripts/prepare_data.py, scripts/relab3.py):
nos_<VID>…, 121 groups) because the trainer packs same-folder clips into one sequence.codec.pth (modded_dac_vq), protobuf shards via build_dataset.py.Community-validated recipe (credit: Enucatl/fish-speech barbero):
r_32_alpha_16_fast (r=32, α=16, α/r=0.5)causal: false (random sampling), max_steps 1200, ckpt every 100, bf16-true, PYTORCH_CUDA_ALLOC_CONF=expandable_segments:Truepython fish_speech/train.py --config-name text2semantic_finetune_egy2 [email protected]_config=r_32_alpha_16_fast⚠️ NEVER LoRA the slow AR with α/r ≥ 1.0: phase-1's first attempt (r32 α32 attention+mlp on both transformers, lr 5e-5) produced pure noise by step 200 — the slow/text transformer collapsed (confirmed by slow/fast ablation, and it is literally a row in the community doc's "didn't work" table).
Required upstream patches (in scripts/patches_notes.md):
llama.py: use_reentrant=False in both checkpoint() calls (LoRA + frozen embeddings otherwise breaks the grad graph)lit_module.py: strict_loading = False on TextToSemantic (checkpoints are LoRA-only; Lightning resume fails otherwise)325-word Egyptian mega-paragraph, masry reference voice:
| System | WER | CER |
|---|---|---|
| s2-pro zero-shot baseline | 0.071 | 0.05 |
| fine-tuned step-1200 (this model) | 0.036–0.06 | 0.007–0.05 |
Fine-tune preserves Egyptian forms the baseline drifts on (e.g. معايا vs معي). Fast-AR tuning = voice/timbre adaptation; text-following stays RL-aligned.
Engine: sglang-omni (S2-Pro supported natively) + OpenAI shim. Measured: TTFA ~0.5 s (streaming, warm radix cache), RTF ~0.85, quality WER 0.00 on server output, output resampled to 24 kHz.
Blackwell/sm_120 porting patches (all required, scripts/patches_notes.md):
site-packages/deep_gemm (asserts on missing CUDA_HOME at import)apt install cuda-nvcc-13-0 libcublas-dev-13-0 libcusparse-dev-13-0 libcusolver-dev-13-0 libcurand-dev-13-0 (torch cu130) + ninjaengine_builder.py: attention backend fa3 → triton (FA3 = Hopper-only)engine_builder.py: disable_cuda_graph: True (graph capture calls FA3 kernels), mem_fraction_static: 0.55audio_decoder.py: FISH_FORCE_SDPA=1 env-gated pure-SDPA replacement for sgl_kernel.flash_attn_with_kvcache (scripts/patch_audio_decoder.py) — Fast-AR attends ≤11 positions, SDPA is exact and fastLaunch:
export CUDA_HOME=/usr/local/cuda-13.0 SGLANG_ENABLE_JIT_DEEPGEMM=0 FISH_FORCE_SDPA=1
sgl-omni serve --model-path <this-repo-dir> --config examples/configs/s2pro_tts.yaml --port 8001
python scripts/serve_shim.py --port 8000 # OpenAI contract: model=s2pro-egy, voices, 24kHz
API (OpenAI-compatible): POST /v1/audio/speech {"model":"s2pro-egy","input":"...","voice":"masry","stream":true,"response_format":"pcm"}.
Voices = voices/<name>.wav+.txt reference pairs (masry = eqkawkab narrator; noselleel/noselleel2/noselleel3 = noselleel narrator candidates).
Goal: deeper Egyptian pronunciation/prosody (slow-AR territory) without breaking RL alignment.
r32 α16 ALL modules (α/r = 0.5, attention+mlp on slow+fast) — the community doc's "minor slow-AR degradation, usable — worth investigating" row. lr 1e-5 cosine, wd 0.01, causal false, max 1200 steps, ckpt every 100.pretrained_ckpt_path → merged dir), so phase-1 voice gains are kept.scripts/serve_eval.py). The failure mode is audio ending early / going quiet / noise — stop immediately if WER degrades, keep phase-1.fishaudio/s2-pro · Training/eval infra: fish-speech main (e5e2926) · Serving: sglang-omni