Downloads · 30 days
0
Yorch233/RSB
RSB is a audio-to-audio model from Yorch233. Use it for the audio-to-audio task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Sep 18, 2026
Parameters
27.8M
111 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors111 MB · 100%
From the Hugging Face model README
Regularized Schrödinger Bridge (RSB) is a generative speech enhancement approach that reconciles fidelity and realism while mitigating exposure bias. RSB regularizes training with a Distortion-Perception Perturbation that constructs time-varying targets by interpolating between clean speech and posterior-mean estimates, and trains the network on perturbed intermediate states to correct toward the ground truth progressively. Consequently, such perturbation simulates inference-time prediction errors, mitigating the training–inference mismatch and thereby reducing exposure bias. Furthermore, it also injects posterior-mean estimates as fidelity-preserving guidance, facilitating reconstruction fidelity.
We have publicly released a checkpoint of generative model, which is based the ncsnpp_base architecture and was trained on the Voicebank+Demand dataset.
There are two ways to download:
CLIpython -m cli.download_pretrained_model
Google Drive.
Download the folder from Google Drive and place it in the pretrained_models/ directory.This project is licensed under the Apache License 2.0.