Downloads · 30 days
0
aoxo/text2asmr-stable-audio
text2asmr-stable-audio is a text-to-audio model from aoxo. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product.
LoRA finetune of stabilityai/stable-audio-open-1.0 on ASMR trigger sounds (tapping, brushing, crinkling, fabric rustling, breathing, etc.), trained to step 2600.
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2026
Repo size
321 MB
Likes
0
Public
Click a slice to open those files.
.ckpt185 MB · 58%
From the Hugging Face model README
LoRA finetune of stabilityai/stable-audio-open-1.0 on ASMR trigger sounds
(tapping, brushing, crinkling, fabric rustling, breathing, etc.), trained to
step 2600.
Status: frozen as v1. Samples at this checkpoint do not reliably produce
clean trigger sounds -- generations exhibit artifacts bleeding in from other
regions of the training distribution (low-frequency hum, snoring-like FX,
mouth clicks/whispered chuckles, broadband hiss) instead of isolated
tapping/brushing. Training on this run has been stopped rather than pushed
further; epoch=0-step=2600.ckpt is kept as the final v1 checkpoint and the
starting point for diagnosing/re-approaching in v2 (likely dataset labeling
or trigger-class separation, not just more steps).
Files:
epoch=0-step=2600.ckpt -- final v1 checkpointv1_sample_tapping.wav, v1_sample_brushing.wav -- raw single-trigger
samples from this checkpoint (audible artifacts described above)v1_composed_demo.wav -- full mixed speech+trigger ASMR composition
(via scripts/compose_asmr.py), same underlying trigger artifacts present
in the non-speech segments