Downloads · 30 days
0
prashantrajbista/deepsc-s
deepsc-s is a audio-to-audio model from prashantrajbista. Use it for the audio-to-audio task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
PyTorch reproduction of DeepSC-S (Weng, Qin & Li, 2021 — arXiv:2012.05369), a neural joint source-channel coding system for speech: instead of separate compression + error-correction + modulation, a single network lea…
Downloads · 30 days
0
Access
Public
Updated Jul 23, 2026
Repo size
1.6 MB
Likes
0
Public
Click a slice to open those files.
.pt1.6 MB · 100%
From the Hugging Face model README
PyTorch reproduction of DeepSC-S (Weng, Qin & Li, 2021 — arXiv:2012.05369), a neural joint source-channel coding system for speech: instead of separate compression + error-correction + modulation, a single network learns to encode speech directly into channel symbols and reconstruct it at the receiver, trained end-to-end through a simulated noisy channel.
Three final-weights checkpoints, each trained with a different simulated channel during
training (ChannelLayer in deepscs/channel.py):
| file | trained on |
|---|---|
deepsc_s_awgn_final.pt | AWGN |
deepsc_s_rayleigh_final.pt | Rayleigh fading |
deepsc_s_rician_final.pt | Rician fading (K=1) |
All three share the same architecture (DeepSC_S in deepscs/model.py):
depth=8 (8 real channel symbols per
frame-column position)Where the paper underspecifies (kernel size, r, D, Rician K, optimizer), the official
TensorFlow repo was used as ground truth
instead of guessing — see docs/initial_plan.md in the code repo for the full list of
flagged choices.
Trained on a 2,000-clip subset of the Edinburgh DataShare / VoiceBank-DEMAND clean 28-speaker set (~10k clips available), 40 epochs. This is a deliberate scope choice for a tractable single-session reproduction, not the paper's full-scale run — expect lower absolute quality (PESQ ~1.0–1.35) than the paper reports. SDR/PESQ curves and the matched-vs-mismatched-channel robustness results are in the write-up linked above.
import torch
from deepscs.model import DeepSC_S # from github.com/prashantrajbista/semantic-communication
model = DeepSC_S(depth=8, n_blocks=4, cardinality=4, r=4)
model.load_state_dict(torch.load("deepsc_s_awgn_final.pt", map_location="cpu"))
model.eval()
# x: (B, 1, F, L) framed audio — see deepscs/audio.py for the framing/deframing utilities
# and deepscs/channel.py for ChannelLayer to simulate the transmission channel.
In: neural transceiver, end-to-end MSE training, AWGN + Rayleigh + Rician, SDR (+ PESQ), cross-channel robustness evaluation.
Out: traditional PCM+Turbo/64-QAM baseline, exact numeric paper reproduction, real RF hardware, estimated-CSI channels, full-scale (~10k clip) training.