Downloads · 30 days
0
nvidia/audio_to_audio_schrodinger_bridge
audio_to_audio_schrodinger_bridge is a machine learning model from nvidia. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Zhifeng Kong, Kevin J Shih, Weili Nie, Arash Vahdat, Sang-gil Lee, Joao Felipe Santos, Ante Jukic, Rafael Valle, Bryan Catanzaro
Downloads · 30 days
0
Access
Public
Updated Aug 1, 2025
Repo size
6.8 GB
Likes
24
Trending 1
Click a slice to open those files.
.ckpt6.8 GB · 100%
From the Hugging Face model README
Zhifeng Kong, Kevin J Shih, Weili Nie, Arash Vahdat, Sang-gil Lee, Joao Felipe Santos, Ante Jukic, Rafael Valle, Bryan Catanzaro
This repo contains the PyTorch implementation of A2SB: Audio-to-Audio Schrodinger Bridges. A2SB is an audio restoration model tailored for high-res music at 44.1kHz. It is capable of both bandwidth extension (predicting high-frequency components) and inpainting (re-generating missing segments). Critically, A2SB is end-to-end without need of a vocoder to predict waveform outputs, and able to restore hour-long audio inputs. A2SB is capable of achieving state-of-the-art bandwidth extension and inpainting quality on several out-of-distribution music test sets.
We propose A2SB, a state-of-the-art, end-to-end, vocoder-free, and multi-task diffusion Schrodinger Bridge model for 44.1kHz high-res music restoration, using an effective factorized audio representation.
A2SB is the first long audio restoration model that could restore hour-long audio without boundary artifacts
The model is provided under the NVIDIA OneWay NonCommercial License.
@article{kong2025a2sb,
title={A2SB: Audio-to-Audio Schrodinger Bridges},
author={Kong, Zhifeng and Shih, Kevin J and Nie, Weili and Vahdat, Arash and Lee, Sang-gil and Santos, Joao Felipe and Jukic, Ante and Valle, Rafael and Catanzaro, Bryan},
journal={arXiv preprint arXiv:2501.11311},
year={2025}
}