Downloads · 30 days
0
SeasonEngine/stable-audio-3-onnx
stable-audio-3-onnx is a machine learning model from SeasonEngine. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
SeasonAudio for Stable Audio 3 Models is a .NET audio generation library for ONNX bundles converted from the Stable Audio 3 Python project.
Downloads · 30 days
0
Access
Public
Updated Aug 29, 2026
Repo size
15.3 GB
Likes
0
Public
Click a slice to open those files.
.onnx_data7.9 GB · 51%
From the Hugging Face model README
SeasonAudio for Stable Audio 3 Models is a .NET audio generation library for ONNX bundles converted from the Stable Audio 3 Python project.
The initial public release focuses on a simple application-facing API:
new StableAudio(model) and Generate(...).wavStableAudio.EnableDebugOutputThe package name is intentionally broader than Stable Audio because the long-term goal is to support more audio model families over time. The current implementation focuses on Stable Audio 3 compatible ONNX workflows, but the project is not an official Stability AI project and does not ship Stability AI model weights.
SeasonAudioSeason.AI.StableAudioIf you prefer to build the ONNX models yourself, clone the repository and run the workflow logic locally.
seedauto, cpu, dml, cuda, coreml, nnapiStableAudio.EnableDebugOutputStableAudio.DebugPromptEncoding(...) and StableAudio.DebugFirstDiTStep(...)dotnet add package SeasonAudio
SeasonAudio currently supports these Stable Audio 3 model ids and aliases:
| Model family | Hugging Face model id | Accepted model values | Typical use |
|---|---|---|---|
| Small SFX | stabilityai/stable-audio-3-small-sfx | stable-audio-3-small-sfx, small-sfx, small_sfx, sa3-sm-sfx | short sound effects, one-shots, UI sounds |
| Small Music | stabilityai/stable-audio-3-small-music | stable-audio-3-small-music, small-music, small_music, sa3-sm-music | lightweight music and loops |
| Medium | stabilityai/stable-audio-3-medium | stable-audio-3-medium, medium, sa3-m | higher-capacity general generation |
You can also pass:
dit.onnxdit.onnx fileBundle layout is expected to resolve the matching decoder and t5gemma tokenizer/text-encoder files alongside the model root.
Model notes:
stable-audio-3-small-sfx is best suited for short clips. For many SFX cases, 3-7 seconds with durationPaddingSeconds around 0-1 gives more stable output.stable-audio-3-small-music is the lightweight choice when you want faster iteration.stable-audio-3-medium is the higher-capacity option when quality matters more than footprint.using Season.AI;
StableAudio.EnableDebugOutput = true; // optional
var stableAudio = new StableAudio(
model: "stable-audio-3-small-sfx",
provider: "cpu");
byte[] wavBytes = stableAudio.Generate(
prompt: "A short sci-fi UI confirmation beep with a soft digital tail",
seconds: 4f,
steps: 8,
cfgScale: 1.5f,
seed: 12345);
File.WriteAllBytes(@"output.wav", wavBytes);
Because the caller environment may not have a predictable audio playback stack, the recommended integration pattern is to save the returned bytes as a .wav file and let the application decide how to play or distribute it.
Example using the music variant:
using Season.AI;
var stableAudio = new StableAudio("stable-audio-3-small-music");
byte[] wavBytes = stableAudio.Generate(
prompt: "Warm ambient synth loop with soft pads and gentle arpeggio",
negativePrompt: "harsh noise, clipping, distorted vocal",
seconds: 12f,
steps: 8,
cfgScale: 2.0f);
File.WriteAllBytes(@"music-preview.wav", wavBytes);
This initial API is intentionally simple:
StableAudio instance with a model alias or a prepared ONNX bundle pathGenerate(...) call returns ready-to-save WAV bytesFor the first release, the recommended runtime target is CPU execution unless you already control the ONNX Runtime provider environment on the target machine.
StableAudio current initialization and generation shape:
public StableAudio(string model, string? provider)
public byte[] Generate(
string prompt,
float seconds = 10f,
int steps = 8,
float cfgScale = 1f,
int? seed = null,
string? negativePrompt = null,
float sigmaMax = 1f,
float apgScale = 1f,
float decoderScale = 1f,
float durationPaddingSeconds = 6f,
bool truncateOutputToDuration = true)
Parameter behavior:
model: constructor parameter; accepts model alias, model bundle directory, or full dit.onnx pathprovider: optional constructor parameter; selects the ONNX Runtime provider during initialization, with supported values auto, cpu, dml, cuda, coreml, nnapiprompt: positive text prompt used for audio generationseconds: requested output duration in seconds; must be greater than 0steps: denoising step count; must be greater than 0cfgScale: classifier-free guidance strength; 1 means no extra guidance branchseed: optional random seed for reproducible outputnegativePrompt: optional negative prompt used when cfgScale is not 1sigmaMax: sampler noise ceiling; valid range is 0.01 to 1.0apgScale: APG guidance factor; valid range is 0 to 1decoderScale: latent scaling before decoder; must be greater than 0durationPaddingSeconds: extra latent duration added before decoding; useful when a model benefits from a little headroomtruncateOutputToDuration: when true, trims the final WAV to seconds; when false, keeps the padded durationPractical tuning guidance:
steps when you want more refinement at the cost of speedcfgScale when prompt adherence is too weakseed when you need deterministic resultsdurationPaddingSeconds for short SFX generation if extra tail content becomes undesirabletruncateOutputToDuration = true for most application-facing output filesEnable debug output before calling the generation API if you need internal sampling diagnostics:
StableAudio.EnableDebugOutput = true;
When enabled:
Debug.WriteLine(...)When disabled:
The default value is:
StableAudio.EnableDebugOutput = false;
Optional troubleshooting helpers:
StableAudio.DebugPromptEncoding(...) returns a text report for tokenizer ids, attention mask, and prompt hidden-state previewsStableAudio.DebugFirstDiTStep(...) returns a deeper text report for the first DiT sampling stepThese helpers are useful for integration validation and prompt/tokenizer diagnostics, but they are not required for normal generation.
StableAudio.Generate(...) returns a byte[] containing a standard .wav file:
44100File.WriteAllBytes(...)This makes the API easy to use in desktop apps, tools, game pipelines, background jobs, and services where the actual playback environment may differ across clients.
Recommended model sources:
Models/...The runtime expects a bundle composed of:
https://stability.ai/licensehttps://huggingface.co/collections/stabilityai/stable-audio-3To export or run your own Stable Audio 3 ONNX bundle, the caller must provide their own Hugging Face credentials and must already have permission to download the source model.
Typical access flow:
HF_TOKEN before running the export script or GitHub workflow.Example:
export HF_TOKEN=hf_your_token_here
PowerShell:
$env:HF_TOKEN = "hf_your_token_here"
If your account does not have access to the selected gated repository, the export script will fail even when HF_TOKEN is present.
The repository includes scripts/export_stable_audio_onnx.py for exporting:
Quick examples:
python scripts/export_stable_audio_onnx.py --variant small-sfx --out-dir out
python scripts/export_stable_audio_onnx.py --variant small-music --out-dir out
python scripts/export_stable_audio_onnx.py --variant medium --out-dir out
The script accepts optional advanced overrides, but --variant is the recommended entry point.
The repository also includes a GitHub Actions workflow:
.github/workflows/export-stable-audio-onnx.ymlThe workflow exposes the same three variants and expects HF_TOKEN to be configured as a repository secret.
HF_TOKEN and export locallySeasonAudio is an independent open source project. It is not affiliated with, endorsed by, or distributed by Stability AI.