Downloads · 30 days
0
imgriff/ce-ssl-ccn2026
ce-ssl-ccn2026 is a feature extraction model from imgriff. Use it when you need embeddings to search or compare text. It is set up for pytorch. The card lists the license as mit.
CochCNN9 checkpoints from the CCN 2026 paper Toward Human-aligned Universal Audio Representations with Contrastive-Equivariant Self-Supervised Learning. This repository hosts the eight models used in the data-scaling…
Downloads · 30 days
0
Access
Public
Updated Aug 27, 2026
Repo size
4.2 GB
Likes
0
Public
Click a slice to open those files.
.safetensors4.2 GB · 100%
From the Hugging Face model README
CochCNN9 checkpoints from the CCN 2026 paper Toward Human-aligned Universal Audio Representations with Contrastive-Equivariant Self-Supervised Learning. This repository hosts the eight models used in the data-scaling experiments: supervised word / auditory-event / multi-task controls, invariant SSL (iSSL), contrastive-equivariant SSL (CE-SSL), and AudioSet-scaled variants of the event and SSL models.
All encoders share the same CochCNN9 backbone and cochleagram frontend (50 ERB filters, 50–10 kHz). Input is a mono waveform at 20 kHz, 2 s clips ((batch, 1, 40000)). Scaled means training on full unbalanced AudioSet; the others use matched Word-Speaker-Noise speech-in-noise (Feather et al., NeurIPS 2019).
| Plot name | Folder | Key | Objective | Data |
|---|---|---|---|---|
| Supervised word | cochcnn9-supervised-word | word | Supervised word classification (794 classes) | Matched Word-Speaker-Noise |
| Supervised auditory events | cochcnn9-supervised-auditory-events | aud_events | Supervised auditory-event classification (517 classes) | Matched Word-Speaker-Noise |
| Supervised multi-task | cochcnn9-supervised-multitask | multitask | Supervised multi-task: word (794), speaker (433), and auditory events (517) | Matched Word-Speaker-Noise |
| Scaled supervised auditory events | cochcnn9-supervised-auditory-events-scaled | scaled_aud_events | Supervised auditory-event classification (527 classes) | AudioSet (scaled) |
| iSSL | cochcnn9-issl | issl | Invariant SSL (Barlow Twins, λ=0) | Matched Word-Speaker-Noise |
| CE-SSL | cochcnn9-ce-ssl | ce_ssl | Contrastive-equivariant SSL (Barlow Twins, λ=0.5) | Matched Word-Speaker-Noise |
| Scaled iSSL | cochcnn9-issl-scaled | scaled_issl | Invariant SSL (Barlow Twins, λ=0) | AudioSet (scaled) |
| Scaled CE-SSL | cochcnn9-ce-ssl-scaled | scaled_ce_ssl | Contrastive-equivariant SSL (Barlow Twins, λ=0.5) | AudioSet (scaled) |
Each folder contains config.yaml (training recipe) and model.safetensors (Lightning state_dict, no optimizer). Model checkpoints should support replication of linear probes, zero-shot evaluations, and brain-model comparisons.
Install the code from GitHub (pip install -e ".[hub]"), then load a checkpoint by registry key:
from lightning_scripts.zero_shot_utils import load_single_cochdnn_model
encoder, name, layer_names = load_single_cochdnn_model(
"ce_ssl", from_hub=True, device="cpu",
)
# waveform: (batch, 1, time) at 20 kHz
activations = encoder(waveform) # dict[str, Tensor], flattened (batch, dim)
Keys: word, aud_events, multitask, scaled_aud_events, issl, ce_ssl, scaled_issl, scaled_ce_ssl. Layers such as relu4 and relufc match the paper evaluations.
To download files directly:
from huggingface_hub import hf_hub_download
config_path = hf_hub_download("imgriff/ce-ssl-ccn2026", "cochcnn9-ce-ssl/config.yaml")
weights_path = hf_hub_download("imgriff/ce-ssl-ccn2026", "cochcnn9-ce-ssl/model.safetensors")
Training code is in the GitHub repository.
Shared settings across these checkpoints: CochCNN9 backbone, LARS optimizer, base learning rate 0.2. Per-model YAML files (epochs, batch size, λ, dataset class) live in each folder as config.yaml and in the GitHub repo under model_configs/. Dataset paths use COCHDNN_* environment variables.
Non-scaled models are trained on matched speech-in-noise from the Word-Speaker-Noise dataset (Feather et al., NeurIPS 2019). Scaled models are trained on full unbalanced AudioSet (Gemmeke et al., ICASSP 2017).
Research on auditory representation learning, including linear probes (ESC-50, Speech Commands, Word-Speaker-Noise word, NSynth), zero-shot triplet evaluations, and brain–model comparisons.
state_dict for exact reconstruction; downstream work typically uses intermediate encoder layers@inproceedings{griffith2026humanaligned,
title={Human-aligned Universal Audio Representations with Contrastive-Equivariant Self-Supervised Learning},
author={Ian M. Griffith and Thomas Edward Yerxa and Josh McDermott and Jenelle Feather},
booktitle={9th Annual Conference on Cognitive Computational Neuroscience},
year={2026},
doi={10.32470/uqprhu8},
url={https://openreview.net/forum?id=qaNtSV4PGm}
}
Word-Speaker-Noise dataset (Feather et al., NeurIPS 2019):
@inproceedings{feather2019metamers,
title={Metamers of neural networks reveal divergence from human perceptual systems},
author={Feather, Jenelle and Durango, Alex and Gonzalez, Ray and McDermott, Josh},
booktitle={Advances in Neural Information Processing Systems},
year={2019}
}
AudioSet (Gemmeke et al., ICASSP 2017):
@inproceedings{gemmeke2017audioset,
title={{Audio Set}: An ontology and human-labeled dataset for audio events},
author={Gemmeke, Jort F. and Ellis, Daniel P. W. and Freedman, Dylan and Jansen, Aren and Lawrence, Wade and Moore, R. Channing and Plakal, Manoj and Ritter, Marvin},
booktitle={Proc. IEEE ICASSP},
year={2017}
}
MIT. See LICENSE in this repository and in the GitHub repo.