Downloads · 30 days
0
ARotting/snip-scope-sae
snip-scope-sae is a machine learning model from ARotting. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
SNIP Scope trains a top-k sparse autoencoder on final-layer hidden activations from the pretrained SNIP-0.4M transformer. A 96-dimensional activation is encoded into a 384-feature overcomplete dictionary, but only the…
Downloads · 30 days
0
Access
Public
Updated Jul 30, 2026
Repo size
298 KB
Likes
0
Public
Click a slice to open those files.
.safetensors297 KB · 94%
From the Hugging Face model README
SNIP Scope trains a top-k sparse autoencoder on final-layer hidden activations from the pretrained SNIP-0.4M transformer. A 96-dimensional activation is encoded into a 384-feature overcomplete dictionary, but only the 16 largest positive features may fire for each token.
The benchmark measures held-out reconstruction, explained variance, active-feature count, dead-feature rate, and token exemplars for each learned feature. A 16-component PCA reconstruction is retained as a dense low-rank control. Feature exemplars are descriptive clues, not proof that a neuron represents one human concept.
The SAE trained on 180,000 final-layer token activations and was measured on 40,000 held-out activations.
| Metric | Top-k SAE | PCA-16 control |
|---|---|---|
| Reconstruction MSE | 0.02585 | 0.29824 |
| Explained variance | 97.41% | 70.11% |
| Active features per token | 16.00 | 16 dense components |
| Dead dictionary features | 3.39% | not applicable |
The learned dictionary contains 384 features and the SAE has 74,112 parameters. The Space exposes each feature's five highest-activating held-out BPE tokens and firing rate. Those exemplars may reflect token identity, syntax, position, or mixed causes; the project does not assign automatic human-readable concepts.
uv run python projects/snip-scope/train.py