Downloads · 30 days
51
27% of all-time downloads
cisco-ai/pase
pase is a audio-to-audio model from cisco-ai. Use it for the audio-to-audio task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
PASE is a state-of-the-art generative speech enhancement model trained to remove noise and reverberation while preserving linguistic content and speaker identity. It operates on 16 kHz mono audio.
Downloads · 30 days
51
27% of all-time downloads
All-time downloads
189
Public
Repo size
1.8 GB
Likes
9
Public
Click a slice to open those files.
.pt1.8 GB · 100%
From the Hugging Face model README
PASE is a state-of-the-art generative speech enhancement model trained to remove noise and reverberation while preserving linguistic content and speaker identity. It operates on 16 kHz mono audio.
PASE contains two main components:
Denoising WavLM (DeWavLM)
Fine‑tuned from WavLM‑Large using denoising representation distillation (DRD).
Performs robust noise supression while effectively mitigating linguistic hallucinations by leveraging the phonological prior from self-supervised WavLM.
Dual‑Stream Vocoder
Reconstructs audio using DeWavLM's dual-stream representations:
Developed by: Copyright © 2026 by Cisco Systems, Inc. All rights reserved.
Cisco product group: Collaboration AI: Xiaobin Rong, Qinwen Hu, Mansur Yesilbursa, Kamil Wojcicki
Model type: Generative Speech Enhancement
License: Apache 2.0
Finetuned from: WavLM-Large
Refer to the repository for quick-start code and examples:
https://github.com/cisco-open/pase
We release a PASE checkpoint that has been trained on an updated list of datasets. For this release, training used:
These source datasets were used to prepare training mixtures and train the released model. The model card and repository do not redistribute the underlying dataset contents; please refer to the original dataset pages and licenses below.
All audio was resampled to 16 kHz.
The performance of the released version compared to the paper's results:
| Model | DNSMOS | UTMOS | SBS | LPS | SpkSim | WER (%) |
|---|---|---|---|---|---|---|
| Vocoder-L24 (paper) | 3.23 | 3.40 | 0.94 | 0.97 | 0.65 | 2.86 |
| Vocoder-L24 (released) | 3.29 | 3.30 | 0.94 | 0.96 | 0.59 | 3.46 |
| DeWavLM (paper) | 3.26 | 3.42 | 0.88 | 0.93 | 0.57 | 7.62 |
| DeWavLM (released) | 3.31 | 3.39 | 0.88 | 0.93 | 0.52 | 7.25 |
| PASE (paper) | 3.12 | 3.09 | 0.90 | 0.93 | 0.80 | 7.49 |
| PASE (released) | 3.08 | 3.21 | 0.91 | 0.94 | 0.80 | 6.76 |
It can be seen that the released version achieves performance very close to that of the paper's results on our simulated test set.
Overall, PASE achieves:
Evaluate outputs for your specific use case. Avoid deployments where misunderstanding enhanced speech could have safety or legal consequences.
If you use PASE in your research, please cite:
@article{PASE,
title={{PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement}},
volume={40},
DOI={10.1609/aaai.v40i39.40562},
number={39},
journal={Proceedings of the AAAI Conference on Artificial Intelligence},
author={Rong, Xiaobin and Hu, Qinwen and Yesilbursa, Mansur and Wojcicki, Kamil and Lu, Jing},
year={2026},
month={Mar.},
pages={32826-32834}
}
Copyright © 2026 by Cisco Systems, Inc. All rights reserved.