Downloads · 30 days
0
suryatmodulus/diamond-1.0
diamond-1.0 is a audio-to-audio model from suryatmodulus. Use it for the audio-to-audio task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Jul 17, 2026
Repo size
673 MB
Likes
0
Public
Click a slice to open those files.
.safetensors667 MB · 99%
From the Hugging Face model README
A sequence-to-sequence model for speech restoration via an autoregressive RQ-Transformer over neural audio codec tokens
Diamond turns degraded audio into near-studio 44.1 kHz speech. A bidirectional Transformer encodes the degraded mel-spectrogram; an autoregressive decoder with cross-attention predicts the tokens of a frozen neural audio codec (Descript Audio Codec, 9-book RVQ), which synthesizes the restored waveform.
Restoration is not denoising. A codec cuts the band above 8 kHz, clipping cuts the peaks, a lost packet zeroes the frame — no mask can return what the observation no longer contains. The missing content has to be generated from the surviving formants, which is why Diamond is a generative seq2seq model rather than a filter.
Diamond is trained from scratch — no pretrained backbone — and reaches the level of the strongest open restorers while remaining the best from-scratch restorer in its class.
Real, heavily degraded speech restored to 44.1 kHz. DNSMOS-P.835 OVRL rises by more than a full point on each clip — audible even without headphones. Try it live in the Diamond Space.
<table> <thead> <tr><th>#</th><th>Degraded input</th><th>Restored — 44.1 kHz</th><th>DNSMOS OVRL</th></tr> </thead> <tbody> <tr> <td>1</td> <td><audio controls src="https://huggingface.co/nineninesix/diamond-1.0/resolve/main/samples/example1_degraded.wav"></audio></td> <td><audio controls src="https://huggingface.co/nineninesix/diamond-1.0/resolve/main/samples/example1_restored.wav"></audio></td> <td>2.16 → <b>3.50</b> (+1.34)</td> </tr> <tr> <td>2</td> <td><audio controls src="https://huggingface.co/nineninesix/diamond-1.0/resolve/main/samples/example2_degraded.wav"></audio></td> <td><audio controls src="https://huggingface.co/nineninesix/diamond-1.0/resolve/main/samples/example2_restored.wav"></audio></td> <td>2.83 → <b>3.59</b> (+0.76)</td> </tr> <tr> <td>3</td> <td><audio controls src="https://huggingface.co/nineninesix/diamond-1.0/resolve/main/samples/example3_degraded.wav"></audio></td> <td><audio controls src="https://huggingface.co/nineninesix/diamond-1.0/resolve/main/samples/example3_restored.wav"></audio></td> <td>2.56 → <b>3.65</b> (+1.10)</td> </tr> <tr> <td>4</td> <td><audio controls src="https://huggingface.co/nineninesix/diamond-1.0/resolve/main/samples/example4_degraded.wav"></audio></td> <td><audio controls src="https://huggingface.co/nineninesix/diamond-1.0/resolve/main/samples/example4_restored.wav"></audio></td> <td>3.32 → <b>3.89</b> (+0.57)</td> </tr> </tbody> </table>Measured on 750 real degraded recordings — identical clips for every model. DNSMOS-P.835 (raw sig_bak_ovr.onnx) for perceptual quality, CER against the ground-truth transcript for content preservation.
| Model | DNSMOS OVRL ↑ | SIG ↑ | BAK ↑ | CER (median) ↓ |
|---|---|---|---|---|
| Sidon | 3.923 | 4.129 | 4.416 | 0.016 |
| Diamond (ours) | 3.829 | 4.056 | 4.348 | 0.028 |
| RE-USE | 3.789 | 4.003 | 4.344 | 0.014 |
| Resemble Enhance | 3.764 | 3.994 | 4.301 | 0.027 |
| UniSE | 3.752 | 4.014 | 4.261 | 0.027 |
| VoiceFixer | 3.566 | 3.790 | 4.226 | 0.042 |
| input (degraded) | 3.575 | 3.890 | 4.082 | 0.014 |
Diamond improves ≈88% of clips (mean ΔOVRL +0.255) and places second of six on perceptual quality — ahead of RE-USE, which trains on roughly four times the data. Both systems that beat it on CER are non-autoregressive: the gap is the exposure bias inherent to AR decoding, not a capacity limit. The one other autoregressive system here, UniSE, lands on exactly Diamond's mean CER (0.131) by a different route — which suggests the tail belongs to the paradigm, not to this particular recipe.
Offline restoration and dataset cleansing — turning large volumes of degraded recordings (podcasts, interviews, archival and user-generated audio that has passed through lossy codecs) into material clean enough to train on. This is the task Diamond was built for and the one it is measured on.
Keep in mind:
Use it responsibly. Restored speech is synthesized speech. Don't present it as an unaltered recording, and don't use it to fabricate or misattribute what someone said.
Built on the Descript Audio Codec for the frozen token space. Trained on LibriTTS-R, Hi-Fi TTS, VCTK. Evaluated with DNSMOS P.835 and faster-whisper.
If you use this work in your research, please cite:
@software{diamond_2026,
author = {Almaz Zholdoshbek uulu, Ulanbek Abdurazakov, Denis Pavlov and Nursultan Bakashov},
title = {Diamond: A Sequence-to-Sequence Model for Speech Restoration via an Autoregressive RQ-Transformer over Neural Audio Codec Tokens},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/nineninesix/diamond-1.0}},
note = {Trained from scratch; no pretrained backbone}
}
@inproceedings{kumar2023dac,
title={High-Fidelity Audio Compression with Improved RVQGAN},
author={Kumar, Rithesh and Seetharaman, Prem and Luebs, Alejandro and Kumar, Ishaan and Kumar, Kundan},
booktitle={NeurIPS},
year={2023},
note={arXiv:2306.06546}
}
@inproceedings{lee2022rqtransformer,
title={Autoregressive Image Generation using Residual Quantization},
author={Lee, Doyup and Kim, Chiheon and Kim, Saehoon and Cho, Minsu and Han, Wook-Shin},
booktitle={CVPR},
year={2022},
note={arXiv:2203.01941}
}
@article{defossez2024moshi,
title={Moshi: a speech-text foundation model for real-time dialogue},
author={D{\'e}fossez, Alexandre and Mazar{\'e}, Laurent and Orsini, Manu and Royer, Am{\'e}lie and P{\'e}rez, Patrick and J{\'e}gou, Herv{\'e} and Grave, Edouard and Zeghidour, Neil},
journal={arXiv preprint arXiv:2410.00037},
year={2024}
}
@inproceedings{copet2023musicgen,
title={Simple and Controllable Music Generation},
author={Copet, Jade and Kreuk, Felix and Gat, Itai and Remez, Tal and Kant, David and Synnaeve, Gabriel and Adi, Yossi and D{\'e}fossez, Alexandre},
booktitle={NeurIPS},
year={2023},
note={arXiv:2306.05284}
}
@inproceedings{koizumi2023librittsr,
title={LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus},
author={Koizumi, Yuma and Zen, Heiga and Karita, Shigeki and Ding, Yifan and Yatabe, Kohei and Morioka, Nobuyuki and Bacchiani, Michiel and Zhang, Yu and Han, Wei and Bapna, Ankur},
booktitle={Interspeech},
year={2023},
note={arXiv:2305.18802}
}
@inproceedings{reddy2022dnsmos,
title={DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors},
author={Reddy, Chandan K. A. and Gopal, Vishak and Cutler, Ross},
booktitle={ICASSP},
year={2022},
note={arXiv:2110.01763}
}
Apache 2.0 — this model and its weights are released under the Apache License 2.0.
The frozen Descript Audio Codec is downloaded at runtime and carries its own license (MIT).