Downloads · 30 days
7
100% of all-time downloads
fernhoof/stable-audio-open-small
stable-audio-open-small is a text-to-audio model from fernhoof. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. The card lists the license as other.
This repository hosts mobile-optimized runtime artifacts and converted weights for Stable Audio Open Small (341M parameters), adapted for on-device generative audio synthesis in the Neural Beats mobile app (iOS & Andr…
Downloads · 30 days
7
100% of all-time downloads
All-time downloads
7
Public
Repo size
8.3 GB
Likes
0
Public
Click a slice to open those files.
.ort1.1 GB · 66%
From the Hugging Face model README
This repository hosts mobile-optimized runtime artifacts and converted weights for Stable Audio Open Small (341M parameters), adapted for on-device generative audio synthesis in the Neural Beats mobile app (iOS & Android, via ONNX Runtime Mobile).
t5_text_encoder.onnx) — encodes the text prompt into cross-attention conditioning embeddings.dit_step.onnx) — a single denoising step of the diffusion transformer; invoked repeatedly (8 rectified-flow pingpong steps) by the on-device sampler.vae_decoder.onnx) — decodes the 64-channel latent produced by the sampler into 44.1 kHz stereo audio (2048× upsampling).All three ONNX models use weight-only INT8 storage: every weight tensor above a small-tensor threshold is symmetrically quantized to int8 (per-output-channel scale for 2D+ weight matrices, per-tensor scale for 1D tensors), with an explicit DequantizeLinear(int8→fp32) op inserted ahead of each tensor's original consumer(s). The compute graph itself is untouched — every op still runs in float32 exactly as before, and both model inputs/outputs remain float32 — so there is no runtime behavior change for native iOS/Android inference code, only a much smaller download (roughly a quarter of the original FP32 size, half of the previous FP16-storage size). Quantization introduces more noise than FP16 storage but remains well within an acceptable range for ambient/generative audio use cases (validated via onnxruntime: no clipping, no NaNs, >0.98 waveform correlation vs. the FP16-storage version on identical inputs).
Each .onnx file is a small graph-structure file; the bulk of each model's weight data lives in a companion .onnx.data external-data file with the same base name (ONNX Runtime resolves this automatically as long as both files are downloaded into the same directory).
| File | Size (approx.) | Description |
|---|---|---|
t5_text_encoder.onnx | ~330 KB | T5-base text encoder graph structure. |
t5_text_encoder.onnx.data | ~105 MB | T5 text encoder weights (INT8 storage). |
dit_step.onnx | ~910 KB | Diffusion transformer single-step graph structure. |
dit_step.onnx.data | ~326 MB | Diffusion transformer weights (INT8 storage). |
vae_decoder.onnx | ~245 KB | Oobleck VAE decoder graph structure. |
vae_decoder.onnx.data | ~75 MB | Oobleck VAE decoder weights (INT8 storage). |
seconds_embedder.json | ~4.2 MB | Tiny seconds_total conditioning embedder weights (FP32 JSON). |
t5_tokenizer_json.json | ~2.3 MB | T5 fast tokenizer definition. |
t5_tokenizer_config.json | ~2.4 KB | T5 tokenizer configuration. |
model_config.json | ~5.5 KB | Complete model architecture hyperparameters & conditioning/sampling configuration. |
model_weights_meta.json | ~0.6 KB | Fast-load metadata summary (sample rate, latent shape, diffusion steps, sampler, component list) for mobile platform channels. |
checksums.sha256 | ~1 KB | SHA-256 checksums for all files above, verified after each file downloads on-device. |
Total download size: ~0.5 GB (down from ~1.0 GB at FP16-storage precision, ~2.0 GB at full FP32 precision).
This model is derived from Stability AI's Stable Audio Open Small and is distributed under the Stability AI Community License.