Downloads · 30 days
5
6% of all-time downloads
rossijakob/toothless-esnet
toothless-esnet is a machine learning model from rossijakob. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
In our paper, we proposed MP-SENet: a TF-domain monaural SE model with parallel magnitude and phase spectra denoising.<br A long-version MP-SENet was extended to the speech denoising, dereverberation, and bandwidth ex…
Downloads · 30 days
5
6% of all-time downloads
All-time downloads
87
Public
Repo size
515 MB
Likes
0
Public
Click a slice to open those files.
.png485 MB · 93%
From the Hugging Face model README
In our paper, we proposed MP-SENet: a TF-domain monaural SE model with parallel magnitude and phase spectra denoising.<br> A long-version MP-SENet was extended to the speech denoising, dereverberation, and bandwidth extension tasks.<br> Audio samples can be found at the demo website.<br> We provide our implementation as open source in this repository.
There is a small bug in our code, but it does not affect the overall performance of the model.
If you intend to retrain the model, it’s strongly recommended to set batch_first=True in the MultiHeadAttention module inside transformer.py, which can significantly reduce the memory usage of the model.
VoiceBank+DEMAND/wavs_clean and VoiceBank+DEMAND/wavs_noisy, respectively. You can also directly download the downsampled 16kHz dataset here.CUDA_VISIBLE_DEVICES=0,1 python train.py --config config.json
Checkpoints and copy of the configuration file are saved in the cp_mpsenet directory by default.<br>
You can change the path by adding --checkpoint_path option.
python inference.py --checkpoint_file [generator checkpoint file path]
You can also use the pretrained best checkpoint files we provide in the best_ckpt directory.
<br>
Generated wav files are saved in generated_files by default.
You can change the path by adding --output_dir option.<br>
Here is an example:
python inference.py --checkpoint_file best_ckpt/g_best_vb --output_dir generated_files/MP-SENet_VB


We referred to HiFiGAN, NSPP and CMGAN to implement this.
@inproceedings{lu2023mp,
title={{toothless-esnet}: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra},
author={hiccup, guru},
booktitle={Proc. Interspeech},
pages={3834--3838},
year={2023}
}
@article{lu2023explicit,
title={Explicit estimation of magnitude and phase spectra in parallel for high-quality speech enhancement},
author={hiccup, guru},
journal={Neural Networks},
volume = {189},
pages = {107562},
year={2025}
}