Downloads · 30 days
0
elliottash/doppelganger
doppelganger is a feature extraction model from elliottash. Use it when you need embeddings to search or compare text. The card lists the license as mit.
Models for the Doppelganger benchmark (matching a synthetic sound effect to the real recording it was generated from). Paper: https://arxiv.org/abs/2607.04337 · Code: https://github.com/elliottash/doppelganger · datas…
Downloads · 30 days
0
Access
Public
Updated Jul 7, 2026
Repo size
918 MB
Likes
0
Public
Click a slice to open those files.
.pt576 MB · 63%
From the Hugging Face model README
Models for the Doppelganger benchmark (matching a synthetic sound effect to the real recording it was generated from). Paper: https://arxiv.org/abs/2607.04337 · Code: https://github.com/elliottash/doppelganger · dataset: https://huggingface.co/datasets/elliottash/doppelganger
heads/ — trained MLP heads (*.head.pt), the paper's models. Each is a compact head
(d → 512 → 256, batch-norm, ReLU, ℓ2-normalized) on top of a frozen encoder. Naming:
<encoder>_ucs_paired[_<variant>]_<objective>.head.pt, with objective ∈ {instance, class} and
per-fold leave-classes-out (kf0..kf4), data-efficiency (sub250..sub4000), and cross-generator
(aldm) variants.
ckpts/ — fine-tuned encoder checkpoints (BEATs, M2D) used in the six-encoder robustness study.import torch
head = torch.load("heads/clap_general_ucs_paired_instance.head.pt", map_location="cpu")
# apply to frozen encoder embeddings (see src/apply_head.py in the code repo)
The heads consume frozen-encoder embeddings; the frozen backbones (CLAP, PANNs, AST, AudioMAE) come
from their original releases, and the BEATs/M2D fine-tunes are in ckpts/.
MIT (heads and fine-tunes). Frozen backbones follow their original licenses.