Downloads · 30 days
31
57% of all-time downloads
Benjonson/s2-pro-fp8
s2-pro-fp8 is a text-to-speech model from Benjonson. Use it when you need text read aloud. It is set up for transformers. The card lists the license as other.
FP8 quantized version of fishaudio/s2-pro.
Downloads · 30 days
31
57% of all-time downloads
All-time downloads
54
Public
Parameters
5.6B
8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6.2 GB · 77%
How the weights are stored.
F8_E4M34B · 72%
From the Hugging Face model README
FP8 quantized version of fishaudio/s2-pro.
Original Model | Technical Report | GitHub | Playground | ComfyUI Node

Fish Audio S2 is an open-sourced text-to-speech system featuring multi-speaker, multi-turn generation, and instruction-following control via natural-language descriptions. The system utilizes a multi-stage training recipe and a staged data pipeline covering video and speech captioning. S2 Pro specifically uses a Dual-Autoregressive (Dual-AR) architecture:
This is a weight-only FP8 quantization of Fish Audio S2 Pro — a state-of-the-art open-source TTS model with fine-grained inline prosody and emotion control across 80+ languages. The quantization cuts the on-disk size roughly in half and reduces VRAM usage from ~24 GB to ~12 GB, with no perceptible quality loss in practice.
| Original (s2-pro) | This (s2-pro-fp8) | |
|---|---|---|
| Weight dtype | bfloat16 | float8_e4m3fn |
| Activation dtype | bfloat16 | bfloat16 |
| Scale | — | per-row float32 |
| File size | ~12 GB | ~6.2 GB |
| VRAM (inference) | ~24 GB | ~12 GB |
| Extra dependencies | none | none |
What is quantized: All nn.Linear weight matrices in both the Slow AR (4B) and Fast AR (400M) backbones — 201 layers in total. Non-linear weights (embeddings, layer norms, codec) remain in bfloat16.
Method: Per-row symmetric FP8
Each output row of every weight matrix has its own float32 scale factor:
scale = max(abs(row)) / FP8_MAX # FP8_MAX = 448.0 for float8_e4m3fn
W_fp8 = round(W_bf16 / scale) # quantize
W_bf16 = W_fp8.to(bfloat16) * scale # dequantize at inference
Per-row scaling captures the per-channel magnitude variation in transformer weight matrices much better than a single per-tensor scale, significantly reducing quantization error at minimal overhead.
No external quantization library required. Dequantization is implemented in pure PyTorch inside a custom FP8Linear module — no torchao, bitsandbytes, or AutoGPTQ needed. The model loads and runs on any machine with PyTorch 2.1+.
File layout inside model.safetensors:
<layer>.weight — float8_e4m3fn tensor (quantized weights)<layer>.weight.scale — float32 tensor, shape [out_features, 1] (per-row scales)bfloat16 (embeddings, norms, codec, etc.)The easiest way to use this model is with ComfyUI-FishAudioS2, which has native support for this FP8 model with zero extra setup.
Install the ComfyUI node via ComfyUI Manager (search FishAudioS2) or manually:
cd ComfyUI/custom_nodes
git clone https://github.com/Saganaki22/ComfyUI-FishAudioS2.git
The model auto-downloads on first use — select s2-pro-fp8 from the model dropdown in any Fish S2 node.
Or download manually:
huggingface-cli download drbaph/s2-pro-fp8 --local-dir ComfyUI/models/fishaudioS2/s2-pro-fp8
precision: auto or bfloat16 — matches the activation dtypeattention: auto or sage_attention for best performancekeep_model_loaded: True if running multiple generations back-to-backWorks with all three nodes: Fish S2 TTS, Fish S2 Voice Clone TTS, Fish S2 Multi-Speaker TTS.
Fish Audio S2 Pro is a leading text-to-speech model with fine-grained inline control of prosody and emotion. Trained on over 10M+ hours of audio data across 80+ languages, it combines reinforcement learning alignment with a Dual-Autoregressive (Dual-AR) architecture.
S2 Pro builds on a decoder-only transformer combined with an RVQ-based audio codec (10 codebooks, ~21 Hz frame rate):
This asymmetric design keeps inference efficient while preserving audio fidelity. The Dual-AR architecture is structurally isomorphic to standard autoregressive LLMs, inheriting LLM-native serving optimizations — continuous batching, paged KV cache, CUDA graph replay, RadixAttention-based prefix caching.
Embed natural-language instructions directly in the text using [tag] syntax. S2 Pro accepts free-form descriptions — not a fixed tag vocabulary:
[pause] [emphasis] [laughing] [inhale] [chuckle] [tsk] [singing] [excited] [volume up] [echo] [angry] [sigh] [whisper] [screaming] [shouting] [surprised] [short pause] [exhale] [delight] [sad] [clearing throat] [shocked] [with strong accent] [professional broadcast tone] [pitch up] [pitch down]
Free-form examples: [whisper in small voice] · [super happy and excited] · [speaking slowly and clearly] · [sarcastic tone]
15,000+ unique tags supported.
Tier 1 (Best Quality): Japanese (ja), English (en), Chinese (zh)
Tier 2: Korean (ko), Spanish (es), Portuguese (pt), Arabic (ar), Russian (ru), French (fr), German (de)
80+ total: sv, it, tr, no, nl, cy, eu, ca, da, gl, ta, hu, fi, pl, et, hi, la, ur, th, vi, jw, bn, yo, sl, cs, sw, nn, he, ms, uk, id, kk, bg, lv, my, tl, sk, ne, fa, af, el, bo, hr, ro, sn, mi, yi, am, be, km, is, az, sd, br, sq, ps, mn, ht, ml, sr, sa, te, ka, bs, pa, lt, kn, si, hy, mr, as, gu, fo
@misc{liao2026fishaudios2technical,
title={Fish Audio S2 Technical Report},
author={Shijia Liao and Yuxuan Wang and Songting Liu and Yifan Cheng and Ruoyi Zhang and Tianyu Li and Shidong Li and Yisheng Zheng and Xingwei Liu and Qingzheng Wang and Zhizhuo Zhou and Jiahua Liu and Xin Chen and Dawei Han},
year={2026},
eprint={2603.08823},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2603.08823},
}
This model inherits the Fish Audio Research License from fishaudio/s2-pro. Research and non-commercial use is permitted free of charge. Commercial use requires a separate license from Fish Audio — contact [email protected].
The FP8 quantization was produced by drbaph and is released under the same license.