Downloads · 30 days
14
52% of all-time downloads
mlr2000/vocoder-small-watermark-detector
vocoder-small-watermark-detector is a feature extraction model from mlr2000. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as cc-by-4.0.
Standalone detector for the fixed 32-bit provenance watermark embedded by the companion VocBulwark vocoder. Given an audio clip, it extracts the watermark bits with the Cage extractor and compares them to the known fi…
Downloads · 30 days
14
52% of all-time downloads
All-time downloads
27
Public
Parameters
650K
2.6 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.6 MB · 99%
From the Hugging Face model README
Standalone detector for the fixed 32-bit provenance watermark embedded by the companion VocBulwark vocoder. Given an audio clip, it extracts the watermark bits with the Cage extractor and compares them to the known fixed code, reporting how many bits match.
Self-contained: loads with trust_remote_code=True, no training repo required.
This detector is part of a set of 6 repositories:
| Repo | Role |
|---|---|
mlr2000/vocoder-large | Large vocoder |
mlr2000/vocoder-large-watermark-detector | Watermark detector for the large model |
mlr2000/vocoder-large-speaker-encoder | Speaker encoder for the large model |
mlr2000/vocoder-small | Small vocoder (generates the watermarked audio) |
mlr2000/vocoder-small-watermark-detector | Watermark detector (this repo) |
mlr2000/vocoder-small-speaker-encoder | Speaker encoder for the small model |
import torchaudio
from transformers import AutoModel
det = AutoModel.from_pretrained("mlr2000/vocoder-small-watermark-detector", trust_remote_code=True).eval()
wav, sr = torchaudio.load("clip.wav") # [C, T]
res = det.detect(wav, input_sample_rate=sr) # resampled to 24 kHz internally
print(res)
# {'matches': 32, 'n_bits': 32, 'p_value': 8.9e-16} # from our model
See example_roundtrip.ipynb in this repo for an end-to-end example
(generate with the vocoder → detect here).
detect() returns a dict with:
matches: how many of the 32 extracted bits equal the fixed code.n_bits: 32.p_value: the probability an unrelated clip matches at least this well under
Binomial(32, 0.5).A clip generated by our vocoder matches all 32 bits (p_value ~ 0); an
unrelated clip matches about half of them (p_value ~ 1). There is no built-in
yes/no threshold — read matches / p_value and pick whatever operating point
you need. Requiring all 32 bits gives an astronomically small false-positive
rate; allowing a few mismatches trades that for robustness to lossy channels.
input_sample_rate, and multi-channel audio is downmixed to mono.mlr2000/vocoder-small. It will not correctly verify audio generated by the large vocoder (mlr2000/vocoder-large), which uses a different fixed 50-bit watermark.If you use this model, please cite:
@misc{muletta2026,
title = {Training a Discriminator-Free Foundation Vocoder
with Integrated Audio Watermarking},
author = {Muletta, Romolo and Deriu, Jan},
year = {2026},
note = {VT2 Project Report, ZHAW School of Engineering}
}
cc-by-4.0. Trained on MLS (CC-BY-4.0) and Common Voice (CC0). Builds
on BigVGAN (MIT) and wav2vec 2.0 (Apache-2.0). Please retain attribution
when redistributing or building on this model.