Downloads · 30 days
0
charlieoneill/embedding-saes
embedding-saes is a machine learning model from charlieoneill. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This repository contains a collection of Sparse Autoencoders (SAEs) trained on embeddings from scientific papers in two domains: Computer Science (cs.LG) and Astrophysics (astro.PH). These SAEs are designed to disenta…
Downloads · 30 days
0
Access
Public
Updated Jul 31, 2024
Repo size
3.5 GB
Likes
16
Public
Click a slice to open those files.
.pth3.5 GB · 100%
From the Hugging Face model README
This repository contains a collection of Sparse Autoencoders (SAEs) trained on embeddings from scientific papers in two domains: Computer Science (cs.LG) and Astrophysics (astro.PH). These SAEs are designed to disentangle semantic concepts in dense embeddings while maintaining semantic fidelity.
The SAEs in this repository are trained on embeddings of scientific paper abstracts from arXiv, specifically from the cs.LG (Computer Science - Machine Learning) and astro.PH (Astrophysics) categories. They are designed to extract interpretable features from dense text embeddings derived from large language models.
Each SAE follows a top-k architecture with varying hyperparameters:
The naming convention for the models is:
{domain}_{k}_{n}_{batch_size}.pth
For example, csLG_128_3072_256.pth represents an SAE trained on cs.LG data with k=128, n=3072, and a batch size of 256.
These SAEs are primarily intended for:
Limitations:
The SAEs were trained on embeddings of abstracts from:
The SAEs were trained using a custom loss function combining reconstruction loss, sparsity constraints, and an auxiliary loss. For detailed training procedures, please refer to our paper (link to be added upon publication).
Performance metrics for various configurations:
| k | n | Domain | MSE | Log FD | Act Mean |
|---|---|---|---|---|---|
| 16 | 3072 | astro.PH | 0.2264 | -2.7204 | 0.1264 |
| 16 | 3072 | cs.LG | 0.2284 | -2.7314 | 0.1332 |
| 64 | 9216 | astro.PH | 0.1182 | -2.4682 | 0.0539 |
| 64 | 9216 | cs.LG | 0.1240 | -2.3536 | 0.0545 |
| 128 | 12288 | astro.PH | 0.0936 | -2.7025 | 0.0399 |
| 128 | 12288 | cs.LG | 0.0942 | -2.0858 | 0.0342 |
For full results, please refer to our paper (link to be added upon publication).
While these models are designed to improve interpretability, users should be aware that:
If you use these models in your research, please cite our paper (citation to be added upon publication).
For more details on the methodology, feature families, and applications in semantic search, please refer to our full paper (link to be added upon publication).