Downloads · 30 days
0
lvjiameng/FAMA-Astro
FAMA-Astro is a feature extraction model from lvjiameng. Use it when you need embeddings to search or compare text. The card lists the license as mit.
FAMA (Foundational Astronomical Masked Autoencoder) is a self-supervised, foundational image model based on the Masked Autoencoder (MAE) architecture, optimized for the unique properties of astronomical data. It is de…
Downloads · 30 days
0
Access
Public
Updated Dec 14, 2025
Repo size
15.7 GB
Likes
0
Public
Click a slice to open those files.
.pth13 GB · 82%
From the Hugging Face model README
FAMA (Foundational Astronomical Masked Autoencoder) is a self-supervised, foundational image model based on the Masked Autoencoder (MAE) architecture, optimized for the unique properties of astronomical data. It is designed to overcome the challenge of heterogeneous, unlabelled image datasets accumulating from wide-field surveys like the DESI Legacy Imaging Surveys and the upcoming Chinese Space Station Telescope (CSST).
The model achieves robust, generalized feature extraction by pre-training the Vision Transformer (ViT) encoder using a high-ratio masking strategy.
FAMA adopts an asymmetric encoder-decoder architecture, utilizing standard ViT models (ViT-B, ViT-L, ViT-H) for the encoder backbone. The lightweight decoder is discarded after pre-training.
| Architecture | Layers | Patch Size | Embed Dim | MLP Size | Heads | Parameters |
|---|---|---|---|---|---|---|
| ViT-Base (FAMA-B) | 12 | 16 | 768 | 3,072 | 12 | 86M |
| ViT-Large (FAMA-L) | 24 | 16 | 1,024 | 4,096 | 16 | 303M |
| ViT-Huge (FAMA-H) | 24 | 14 | 1,536 | 6,144 | 16 | 680M |
The weights provided below are the pre-trained encoders (ViT-B, ViT-L, ViT-H) from the self-supervised MAE phase, ready for transfer learning via fine-tuning or linear probing.
| Model Size | Weights File | Pre-train Data |
|---|---|---|
| FAMA-B | base_patch16.pth | DESI-1M |
| FAMA-L | large_patch16.pth | DESI-1M |
| FAMA-H | huge_patch14.pth | DESI-1M |
Note: The DESI-1M dataset was a random sample of 2 million galaxies from the DESI Legacy Imaging Surveys DR9, augmented with background cutouts. The actual DESI-1M dataset size is 1 million samples used for the pre-training experiment.
FAMA models were rigorously validated across three distinct transfer learning tasks: Classification (Galaxy Morphology), Regression (Photometric Redshift), and Detection (Gravitational Lensing).
| Method | backbone | Pre-train Data | Acc on galaxy-desi | Acc on galaxy-sdss |
|---|---|---|---|---|
| FAMA (ours) | ViT-H | DESI-1M | 89.10 | 96.02 |
FAMA achieves the highest Average Precision (AP) scores for strong gravitational lensing detection using the ViTDet adaptation.
| Method | backbone | AP | AP<sup>75</sup> |
|---|---|---|---|
| FAMA (ours) | ViT-H | 42.62 | 49.43 |
The pre-trained model on DESI data is fine-tuned on the SDSS Redshift dataset.
| backbone | Δz (Bias, Lower is Better) | σ<sub>MAD</sub> (Dispersion, Lower is Better) |
|---|---|---|
| FAMA ViT-H | 0.51 × 10⁻⁴ | 0.56 × 10⁻² |
The following steps outline the use of the FAMA encoder weights for fine-tuning on a downstream task (e.g., classification).
The model was pre-trained using the following data processing steps:
Load the weights into a standard ViT encoder and attach a task-specific head.
The following fine-tuning configurations were used for the galaxy-desi classification task:
| Config | ViT-Base | ViT-Large | ViT-Huge |
|---|---|---|---|
| Optimizer | AdamW | AdamW | AdamW |
| Learning Rate | 1.5 × 10⁻³ | 2 × 10⁻³ | 1 × 10⁻³ |
| Batch Size | 64 | 64 | 32 |
| TrainingEpochs | 50 | 50 | 50 |
| LR Schedule | Cosine Decay | Cosine Decay | Cosine Decay |
If you use FAMA in your research, please cite the associated work:
@article{FAMA_2025,
title={FAMA -- a Scalable Foundational Astronomical Masked Autoencoder for Astronomical Image Analysis},
author={Lv, Jiameng and Li, Xu and Cao, Liang and Gao, Xi and Li, Nan and Fu, Mingxiang and Li, Yushan and Duan, Manni and Jia, Peng},
journal={Preprint submitted to Elsevier},
year={2025}
}