Downloads · 30 days
3
13% of all-time downloads
ryanontheinside/stable-audio-3-optimized-fp8
stable-audio-3-optimized-fp8 is a machine learning model from ryanontheinside. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for tensorrt. The card lists the license as other.
FP8 GEMM-trunk quantization of the Stable Audio 3 medium DiT, built from stabilityai/stable-audio-3-optimized onnx/sa3-m/ditfp16mixed.onnx with the producer recipe in Stability-AI/stable-audio-3 PR 47 (build/makecalib…
Downloads · 30 days
3
13% of all-time downloads
All-time downloads
24
Public
Repo size
5.9 GB
Likes
0
Public
Click a slice to open those files.
.data2.9 GB · 66%
From the Hugging Face model README
FP8 GEMM-trunk quantization of the Stable Audio 3 medium DiT, built from
stabilityai/stable-audio-3-optimized onnx/sa3-m/dit_fp16mixed.onnx with the
producer recipe in Stability-AI/stable-audio-3 PR #47
(build/make_calib.py + build/build_dit_fp8.py). This is a derivative of
Stability AI's model weights and is distributed under the Stability AI
Community License; see the base model for terms.
onnx/sa3-m/dit_fp8.onnx + dit_fp8.onnx.data — the quantized ONNX
(arch-independent; compile with build_from_onnx.py sa3-m-fp8, plain
STRONGLY_TYPED, no ModelOpt needed)tensorRT/sm_120/sa3-m/dit_fp8.trt — prebuilt engine for RTX 50xx
(sm_120), TensorRT 10.16.1.11. TRT engines are not portable across GPU
architectures or TRT minor versions; rebuild from the ONNX for anything else.Inputs/outputs are FP32, drop-in for the FP16-mixed DiT engine
(sa3_trt --precision fp8, paired with the FP16-mixed decoder).