Downloads · 30 days
0
sammoran-phd/cara-native-stable-audio
cara-native-stable-audio is a text-to-audio model from sammoran-phd. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. It is set up for stable-audio-tools. The card lists the license as other.
This is the Phase 2 CARA-native Stable Audio Open Small checkpoint used in the CARA cross-architecture attribution research. The fork exposes batch-aligned native DiT features and trains a checkpoint-owned hierarchica…
Downloads · 30 days
0
Access
Public
Updated Jul 23, 2026
Repo size
1.7 GB
Likes
0
Public
Click a slice to open those files.
.ckpt1.7 GB · 97%
From the Hugging Face model README
This is the Phase 2 CARA-native Stable Audio Open Small checkpoint used in the CARA cross-architecture attribution research. The fork exposes batch-aligned native DiT features and trains a checkpoint-owned hierarchical 98-pool/9-family head together with the diffusion model.
The release is an open-weight peer-review artifact, not an OSI open-source model. Its use and redistribution remain subject to the Stability AI Community License.
phase2/native_stable_audio_full.ckpt: completed 7,665-step model/head
checkpoint.phase2/native_stable_audio_training_report.json: training contract.registry/: exact CARA pool/family ordering used by the head.evidence/: authoritative prompt-visible and fixed-audio score reports.source/: exact source snapshot used to load and evaluate the checkpoint.LICENSE and NOTICE: required upstream license and attribution.cara_model_manifest.json: byte sizes and SHA-256 values for release files.Use the immutable phase2-v1 tag, or the Hub commit hash it resolves to.
hf download sammoran-phd/cara-native-stable-audio \
--revision phase2-v1 \
--local-dir cara-native-stable-audio-release
mkdir cara-native-stable-audio-source
tar -xzf \
cara-native-stable-audio-release/source/cara-native-stable-audio-source.tar.gz \
-C cara-native-stable-audio-source
cd cara-native-stable-audio-source
Load stabilityai/stable-audio-open-small with stable-audio-tools, attach the
included fork's CARAAttributionHead, and load
phase2/native_stable_audio_full.ckpt. The included benchmark scripts perform
the registry-hash, global-step, native-head, and feature-shape checks before
evaluation. The complete invocation is in the cara-native-musicmodels Phase 2
job specification.
On the primary 780-waveform balanced fixed-audio core, this checkpoint scored 7.82% exact top-1, 25.64% top-3, 33.21% pool-derived family accuracy, 6.90% ECE, and 100% registry-valid output. On the earlier 320-row prompt-visible benchmark it scored 23.75% exact top-1, 56.56% top-3, and 99.38% family accuracy.
The fixed-audio result is the primary estimate of audio-conditioned attribution. The much higher prompt-visible family score is dominated by visible semantic taxonomy and is not evidence of reliable exact source identification.
This release is intended for academic reproduction, interface auditing, and controlled attribution experiments. It is not a provenance, copyright identification, royalty allocation, or safety system. The primary fixed-audio core covers 39 of 98 pools and represents one source corpus and one training run.
The stable-audio-tools code is MIT licensed. The base model and this derivative
checkpoint are governed by the Stability AI Community License. Redistribution
must include that agreement and the required NOTICE; commercial conditions
depend on the user's circumstances and the current upstream license.