Downloads · 30 days
0
AntonyT1207/AstroCo
AstroCo is a feature extraction model from AntonyT1207. Use it when you need embeddings to search or compare text. It is set up for pytorch. The card lists the license as mit.
Self-supervised representation learning for irregular, sparsely-sampled astronomical light curves.
Downloads · 30 days
0
Access
Public
Updated Jun 19, 2026
Repo size
282 MB
Likes
1
Public
Click a slice to open those files.
.ckpt254 MB · 87%
From the Hugging Face model README
Self-supervised representation learning for irregular, sparsely-sampled astronomical light curves.
AstroCo pretrains a Conformer-style encoder on raw MACHO R-band light curves with a masked-reconstruction objective, then transfers the frozen embedding to downstream variable-star classification with very few labels. It improves reconstruction error by 61 to 70% over the Astromer baselines and sets a stronger few-shot transfer point on the Alcock benchmark.
First-author work, NeurIPS 2025 ML4PS workshop.
| Model | Layers | Width (d) | Params | Recon RMSE | R² | Pretrain GPU |
|---|---|---|---|---|---|---|
astroco_s.ckpt | 4 | 276 (4 heads x 69) | 5.9M | 0.060 | 0.922 | A100 80GB |
astroco_l.ckpt | 12 | 256 (4 heads x 64) | 15.2M | 0.044 | 0.956 | H200 |
Both are PyTorch Lightning checkpoints. RMSE is masked-reconstruction error on held-out MACHO R-band; lower is better.
| Model | RMSE |
|---|---|
| Astromer v1 | 0.148 |
| Astromer v2 | 0.113 |
| AstroCo-S | 0.060 |
| AstroCo-L | 0.044 |
AstroCo-L is 70% below Astromer v1 and 61% below Astromer v2.
The encoder is frozen after pretraining; only a linear probe is trained on a small number of labels per class. Scores are 3-fold averages.
| Labels / class | AstroCo-S | AstroCo-L |
|---|---|---|
| 20 | 66.61 | 67.57 |
| 100 | 74.85 | 75.88 |
| 500 | 79.10 | 79.23 |
The gain holds in the low-label regime, which is where a transferable representation matters most.
A Conformer-style encoder built for irregular time series. Each block stacks:
Positional information uses an Astromer-style embedding so the model reads irregular sampling directly. Pretraining is masked reconstruction: 50% of points probed, 60% masked, with a learned mask token. Inputs are 200-point windows, brightness and time zero-mean normalized, trained at fp16 with DDP.
Encoder block source: astro_model_arch/Astroco.py. Full hyperparameters: astro_model_arch/hyparams_astroco_{s,l}.yaml.
import torch, yaml
from astro_model_arch.Astroco import Astroco # encoder definition in this repo
cfg = yaml.safe_load(open("astro_model_arch/hyparams_astroco_l.yaml"))
model = Astroco(**cfg) # the yaml keys are the constructor kwargs
ckpt = torch.load("astroco_l.ckpt", map_location="cpu")
model.load_state_dict(ckpt.get("state_dict", ckpt), strict=False)
model.eval()
Use hyparams_astroco_s.yaml with astroco_s.ckpt. The checkpoints carry the Lightning training state, so strict=False skips the loss and mask-token buffers when you only want the encoder.
The model reads one nested dict. Each tensor is shaped (batch, window, 1) with window = 200, single band (_0, MACHO R). Brightness and time are zero-mean normalized per window; errors are the photometric uncertainties.
B, L = 4, 200
batch = {"LC": {
"brightness_0": torch.randn(B, L, 1), # magnitudes, zero-mean normalized
"time_0": torch.randn(B, L, 1), # observation times, zero-mean normalized
"brightness_err_0": torch.randn(B, L, 1), # photometric errors
}}
emb, _ = model(batch)
z = emb["LC"]["z_emb_12"] # (B, 200, 256): layer-weighted embedding from AstroCo-L
The forward pass returns per-layer embeddings under z_emb_0 .. z_emb_N; the last index holds the softmax-weighted sum across layers and is the one to use. For AstroCo-L that key is z_emb_12 (width 256); for AstroCo-S it is z_emb_4 (width 276). Pool over the time axis (mean, or a CLS-style aggregate) to get a per-light-curve vector for a downstream head.
time_0 (MJD, in days), brightness_0 (R-band magnitude), and brightness_err_0 (photometric error), plus an ID. Length is variable, a few hundred points per curve. Stored as WebDataset .tar.gz shards (one .pth per array), split into train / val / test folds.(batch, 200, 1) dict the forward pass expects.classification_data_link.md.@inproceedings{tan2025astroco,
title = {AstroCo: Self-Supervised Representation Learning for Irregular Astronomical Light Curves},
author = {Tan, Antony},
booktitle = {NeurIPS 2025 Workshop on Machine Learning and the Physical Sciences (ML4PS)},
year = {2025}
}
classification_data_link.mdastro_model_arch/astroco_results/test_results/