Downloads · 30 days
0
Gautam0901/chola-compressor
chola-compressor is a machine learning model from Gautam0901. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
A compact spectrogram U-Net (1.93M parameters) that restores speech quality lost to Codec2 compression, trained via knowledge distillation for use in off-grid LoRa voice communication links. Designed to run in real ti…
Downloads · 30 days
0
Access
Public
Updated Aug 16, 2026
Repo size
8.4 MB
Likes
1
Public
Click a slice to open those files.
.pt7.8 MB · 90%
From the Hugging Face model README
A compact spectrogram U-Net (1.93M parameters) that restores speech quality lost to Codec2 compression, trained via knowledge distillation for use in off-grid LoRa voice communication links. Designed to run in real time on edge hardware (Jetson-class), not in the cloud.
LoRa radio has enough bandwidth for Codec2 at very low bitrates (this project uses the 1200bps mode), which makes off-grid voice communication possible but introduces heavy compression artifacts. This model sits after Codec2 decoding on the receiving end and restores some of that lost quality.
3-level U-Net operating on STFT magnitude spectrograms (n_fft=512, hop=128, 16kHz):
[0,1] mask multiplied against the input magnitude (not raw magnitude directly — more stable to train)models_no_teacher) was trained with teacher_weight=0, i.e. supervised directly against real clean speech.log1p), chosen after finding raw-magnitude L1 over-weights loud regions and under-penalizes quiet, perceptually important detailEvaluated on a held-out, speaker-disjoint test split (403 utterances). All three metrics computed against the true clean reference.
| Candidate | PESQ ↑ | STOI ↑ | SI-SDR (dB) ↑ |
|---|---|---|---|
| Codec2-degraded (no processing) | 1.522 | 0.658 | -28.23 |
| Denoiser (teacher, for reference) | 1.543 | 0.650 | -28.08 |
| This model | 1.471 | 0.780 | -26.69 |


Real, verified gains: +0.12 STOI (intelligibility) and +1.5dB SI-SDR over doing nothing. Both metrics are dominated by energy/envelope accuracy, where this model clearly helps.
Known limitation — PESQ: PESQ is highly sensitive to phase accuracy, and this model only predicts a magnitude mask, reconstructing with the degraded audio's original (uncontrolled) phase. That's the most likely explanation for PESQ landing slightly below the unprocessed baseline despite STOI/SI-SDR improving substantially — three independent loss-function reformulations (distillation weight, log vs. linear magnitude) all left PESQ in the same 1.47–1.48 range, which is consistent with a phase-reconstruction ceiling rather than a loss-tuning problem. A phase-aware architecture (predicting complex spectrograms or a phase correction term) would likely be needed to close this gap; that's a known next step, not yet implemented in this checkpoint.

Research and portfolio demonstration of magnitude-domain speech restoration via knowledge distillation for bandwidth-constrained voice links. Not validated for safety-critical or emergency-communication deployment.
import torch
import numpy as np
import librosa
import soundfile as sf
N_FFT, HOP_LENGTH, SAMPLE_RATE = 512, 128, 16000
class SpectrogramUNet(torch.nn.Module):
# ... see model.py in the project repo for the full class definition
pass
def restore(degraded_wav_path, model, device="cpu"):
audio, sr = sf.read(degraded_wav_path, dtype="float32")
if sr != SAMPLE_RATE:
audio = librosa.resample(audio, orig_sr=sr, target_sr=SAMPLE_RATE)
stft = librosa.stft(audio, n_fft=N_FFT, hop_length=HOP_LENGTH)
mag, phase = np.abs(stft), np.angle(stft)
x = torch.from_numpy(mag).float().unsqueeze(0).unsqueeze(0).to(device)
with torch.no_grad():
pred_mag = model(x).squeeze().cpu().numpy()
restored_stft = pred_mag * np.exp(1j * phase) # reuses degraded audio's phase
return librosa.istft(restored_stft, hop_length=HOP_LENGTH)
model = SpectrogramUNet(base_channels=32)
model.load_state_dict(torch.load("best.pt", map_location="cpu"))
model.eval()
restored = restore("degraded.wav", model)
sf.write("restored.wav", restored, SAMPLE_RATE)
Full training/eval/live-test code: see the project repository.
If you use this model, please cite the LoRaVoiceLink project (link to source repo).