Downloads · 30 days
9
47% of all-time downloads
TechnoBaptist/stupase
stupase is a audio-to-audio model from TechnoBaptist. Use it for the audio-to-audio task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
StuPASE is a state-of-the-art generative speech enhancement model trained to remove noise and reverberation while preserving linguistic content and speaker identity, and achieving studio-level perceptual quality. It o…
Downloads · 30 days
9
47% of all-time downloads
All-time downloads
19
Public
Repo size
2.3 GB
Likes
0
Public
Click a slice to open those files.
.pt2.3 GB · 100%
From the Hugging Face model README
StuPASE is a state-of-the-art generative speech enhancement model trained to remove noise and reverberation while preserving linguistic content and speaker identity, and achieving studio-level perceptual quality. It operates on 16 kHz mono audio.
StuPASE contains three main components:
DeWavLM-R: Performs low-hallucination phonetic enhancement, fine‑tuned from DeWavLM using dry targets for improved dereverberation.
CFM: Performs phonetic-guided acoustic enhancement.
Mel Vocoder: Reconstructs enhanced wavforms.
Developed by: Copyright © 2026 by Cisco Systems, Inc. All rights reserved.
Cisco product group: Collaboration AI: Xiaobin Rong, Mansur Yesilbursa, Kamil Wojcicki
Model type: Generative Speech Enhancement
License: Apache 2.0
Finetuned from: WavLM-Large, DeWavLM
Refer to the repository for quick-start code and examples:
https://github.com/cisco-open/pase
We release a StuPASE checkpoint that has been trained on an updated list of datasets. For this release, training used:
These source datasets were used to prepare training mixtures and train the released model. The model card and repository do not redistribute the underlying dataset contents; please refer to the original dataset pages and licenses below.
All audio was resampled to 16 kHz.
The performance of the retrained version compared to the original one:
| Model | DNSMOS | UTMOS | SBS | LPS | SpkSim | WER (%) |
|---|---|---|---|---|---|---|
| DeWavLM-R (orig.) | 3.35 | 3.94 | 0.84 | 0.88 | 0.49 | 13.22 |
| DeWavLM-R (retrained) | 3.38 | 3.62 | 0.85 | 0.89 | 0.41 | 12.27 |
| StuPASE (orig.) | 3.37 | 4.08 | 0.85 | 0.90 | 0.68 | 11.57 |
| StuPASE (retrained) | 3.36 | 4.02 | 0.85 | 0.89 | 0.66 | 12.06 |
It can be seen that the retrained version achieves performance very close to that of the original version on our simulated test set.
Overall, StuPASE achieves:
Evaluate outputs for your specific use case. Avoid deployments where misunderstanding enhanced speech could have safety or legal consequences.
If you use StuPASE in your research, please cite:
@misc{StuPASE,
title={{StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement}},
author={Xiaobin Rong and Jun Gao and Zheng Wang and Mansur Yesilbursa and Kamil Wojcicki and Jing Lu},
year={2026},
eprint={2603.09234},
archivePrefix={arXiv},
primaryClass={eess.AS},
url={https://arxiv.org/abs/2603.09234},
}
Copyright © 2026 by Cisco Systems, Inc. All rights reserved.