Downloads · 30 days
0
bovod-sjtu/Audio-align
Audio-align is a machine learning model from bovod-sjtu. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as apache-2.0.
The auditable paper sequence is pinned at revision a9324d03830d58fca87eac010162982efa106b22. It contains four fp16 active-interface checkpoints:
Downloads · 30 days
0
Access
Public
Updated Aug 25, 2026
Repo size
82.1 GB
Likes
0
Public
Click a slice to open those files.
.pt74.2 GB · 90%
From the Hugging Face model README
The auditable paper sequence is pinned at revision
a9324d03830d58fca87eac010162982efa106b22.
It contains four fp16 active-interface checkpoints:
| Directory | Role |
|---|---|
aligned-stage3-step45000 | Aligned real-audio base |
pre-zero-audio-phase2-step260000 | Direct before-checkpoint |
zero-audio-phase2.2-step24000 | About 520K pure-text QA; no waveforms |
zero-audio-phase2.3-step5000 | About 120K pure-text chat; no waveforms |
Each directory contains TASU, the acoustic projector, and the instruction router: 84,655,552 parameters, or 5.4839% of the main LLM's 1,543,714,304 independent parameters. The tied token embedding and language-model head are counted once rather than as separate state-dict aliases. The accompanying verification records exact tensor equality for Qwen2.5-1.5B-Instruct, SenseVoice, and ArCap across the sequence.
The claim is deliberately narrower than “training without audio”: real audio is used for grounding before the two zero-audio post-alignment stages. Current evidence is strongest for semantic and speech-paralinguistic transfer; general non-speech signal/event transfer still needs a matched causal experiment.
AIR-Bench metrics in the manifest distinguish historical valid-only accuracy from coverage-aware end-to-end accuracy; unparsable responses are incorrect in the latter.
See the release manifest, frozen-state verification, and GitHub documentation.
The repository root also stores the original full non-main-LLM export for
phase2_3_chat/model-5000.pt. It includes the frozen SenseVoice and ArCap
experts as well as the Phase 2.3 interfaces:
| File | Contents |
|---|---|
semantic_branch.safetensors | SenseVoice, legacy semantic layers, TASU |
acoustic_branch.safetensors | ArCap and acoustic projector |
router.safetensors | BERT-small instruction router |
manifest.json | Tensor counts, provenance, and SHA256 checksums |
These frozen expert tensors can reconstruct every checkpoint in the paper timeline because the exact comparison proves that they never change. Load the full-branch state first, then overwrite the three active interfaces with the selected timeline directory. Qwen2.5-1.5B-Instruct is loaded separately from its upstream release.
For an unwrapped PyTorch model, strip the training-time module. prefix:
from safetensors.torch import load_file
state = {}
for filename in [
"semantic_branch.safetensors",
"acoustic_branch.safetensors",
"router.safetensors",
]:
state.update({
key.removeprefix("module."): value
for key, value in load_file(filename, device="cpu").items()
})
missing, unexpected = model.load_state_dict(state, strict=False)
assert not unexpected
The historical Phase 2.3 checkpoint reports ASR test-clean/test-other WER 3.50%/7.22%, AIR-Bench Chat 2.365/10 with a Qwen3.5-27B judge, and VoiceBench sd-qa 21.8%. Metrics, failure cases, and claim boundaries are documented in the GitHub repository rather than inferred from a single aggregate score.