Downloads Β· 30 days
0
willychap/camulator
camulator is a other model from willychap. Use it for the other task on the model card, and read the license before you ship it in a product. It is set up for miles-credit. The card lists the license as apache-2.0.
CAMulator is an auto-regressive machine-learned emulator of NSF NCAR's CAM6 atmosphere, trained and run within the CREDIT framework. Given prescribed sea-surface temperature, sea-ice, incoming solar radiation, and CO2β¦
Downloads Β· 30 days
0
Access
Public
Updated Sep 16, 2026
Repo size
84.7 GB
Likes
0
Public
Click a slice to open those files.
.pt71.5 GB Β· 84%
From the Hugging Face model README
CAMulator is an auto-regressive machine-learned emulator of NSF NCAR's CAM6 atmosphere, trained and run within the CREDIT framework. Given prescribed sea-surface temperature, sea-ice, incoming solar radiation, and CO2, it rolls a 1 degree (192x288), 32-level, 6-hourly atmospheric state forward for climate-length simulations (years to decades). It conserves global dry-air mass, moisture, and total atmospheric energy, remains numerically stable over decadal rollouts, and reproduces the annual climatology together with major modes of variability such as ENSO and the NAO -- at roughly a 350x speedup over CAM6, making it an efficient way to generate large climate ensembles.
The model and method are described in Chapman et al. (2025), CAMulator: Fast Emulation of the Community Atmosphere Model (arXiv:2504.06007).
CAMulator is a research tool. It emulates a specific CAM6 configuration and is not a substitute for an operational forecast or a full Earth-system model.
camulator_huggingface, dir climate/)# 1. get the toolbox
git clone -b camulator_huggingface https://github.com/WillyChap/miles-credit.git camulator
cd camulator
# 2. environment (PyTorch 2.4.1 + CUDA 12.1; pinned in environment.yml)
conda env create -f environment.yml -n camulator
conda activate camulator
pip install -e . --no-deps # --no-deps: environment.yml pins the exact stack
# 3. pull this model + its inputs into ./assets/
cd climate
python download_assets.py --repo_id willychap/camulator # default checkpoint = epoch 69
# 4. verify + run (writes monthly-mean NetCDF by default)
python check_setup.py
bash RunQuickClimate.sh
Full instructions, configuration, and the asset manifest are in
climate/README.md.
We evaluate many training checkpoints by running each as a free-running, autoregressive 35-year rollout (1980-2014, 6-hourly, no-leap) and scoring it against the CREDIT ERA5-scaled training target on the same 1 degree grid (latitude-weighted), looking at both the monthly-mean climatology and the 6-hourly distribution.
Training loss is a poor guide to whether an emulator will hold together for decades, so we select checkpoints on emulator behaviour, not on validation loss. Every candidate epoch gets the same treatment:
score_TREFHT and score_PRECT; the combined score is their sum. Lower is
better. A separate average-rank ranking is computed as an independent
cross-check that the result is not an artefact of z-scoring.This matters because the ranking is not what a loss curve would suggest, and the metrics disagree with each other β see the trade-off discussion below.
All evaluation rollouts are run with the global energy fixer active, matching the constraint the model was trained under (see Conservation at inference below).
Checkpoint 69 is selected as the default. The rest of the top ten by combined
score (epochs 36, 70, 66, 51, 47, 75, 52, 64, 65) are provided as well, so you can
evaluate them yourself, build cheap checkpoint ensembles, or study sensitivity to
training stage β pick one with
download_assets.py --checkpoint checkpoint.pt000NN.pt. Several checkpoints from
earlier selections (33, 43, 46, 48, 63, 68, 76) also remain hosted.
Checkpoint 69 climatology (latitude-weighted, full 35-yr record):
| Field | Spatiotemporal RMSE | Global-mean bias | Decadal-trend error | Global-mean monthly RMSE | Annual corr. |
|---|---|---|---|---|---|
| TREFHT | 1.602 K | -0.078 K | +0.013 K/decade | 0.160 K | 0.989 |
| PRECT | 6.12e-4 (2.4 mm/day) | +3.2e-6 (+0.01 mm/day) | -- | clim. RMSE 7.78e-5 | -- |
Checkpoint 69 is the default because it ranks first on precipitation overall and first on the blended monthly + 6-hourly score. On the monthly-mean climatology alone it is effectively tied with checkpoint 36 (the two composite scores differ by 1%, and the independent average-rank cross-check mildly prefers 36); the 6-hourly tails are what separate them, and there 69 wins decisively. But the choice still encodes a weighting, and it is worth being explicit about the trade-off, because no single checkpoint wins everything:
So: use 69 for general-purpose work and anything precipitation-driven; consider 70 if you are temperature-focused, or 36 if you want the best monthly-mean temperature and do not care about rainfall extremes. The default reflects a 50/50 blend of the monthly and 6-hourly scores with temperature and precipitation weighted roughly equally; a different, equally defensible weighting would select a different checkpoint.

This is the most informative view of the selection, and it is worth reading carefully before you assume "later epoch = better model".
The remaining figures summarize the outcome: a skill scorecard across the top checkpoints, checkpoint 69's annual-mean bias maps, the global-mean temperature evolution against the training target, and the 6-hourly precipitation distribution. The 6-hourly tail scores that break ties between the finalists follow in their own subsection.




Note on PRECT units: native values are metres of liquid-water equivalent per 6-hourly step (ERA5
tpconvention); mm/day = native x 4000 (1000 mm/m x 4 steps/day). Checkpoint 69's global-mean precipitation is 2.95 mm/day vs. truth 2.93.
The five finalists from the monthly sweep (checkpoints 69, 70, 52, 36, 46) are then scored on the tails of the 6-hourly TREFHT and PRECT distributions over the full 35-year record. Each tail statistic is computed at every grid point, and the model map is compared with the truth map using the error measures defined in the key below. PRECT statistics are given in mm/day (native x 4000).


Bias-RMSE of each tail statistic against truth (lower is better; best per row in bold):
| Tail statistic | Truth global mean | ckpt 69 | ckpt 70 | ckpt 52 | ckpt 36 | ckpt 46 |
|---|---|---|---|---|---|---|
| TREFHT p1 (cold) | 278.074 K | 0.939 | 0.915 | 1.132 | 1.143 | 1.196 |
| TREFHT p0.1 (cold) | 275.932 K | 1.240 | 1.257 | 1.467 | 1.411 | 1.444 |
| TREFHT p99 (hot) | 296.069 K | 0.702 | 0.682 | 0.651 | 0.686 | 0.697 |
| TREFHT p99.9 (hot) | 297.535 K | 0.903 | 0.872 | 0.886 | 0.881 | 0.876 |
| TREFHT TXx | 297.140 K | 0.935 | 0.866 | 0.883 | 0.871 | 0.866 |
| TREFHT TNn | 276.765 K | 1.057 | 0.995 | 1.182 | 1.201 | 1.251 |
| PRECT p99 | 31.78 mm/day | 3.56 | 3.78 | 3.53 | 4.17 | 3.76 |
| PRECT p99.9 | 65.16 mm/day | 9.52 | 10.24 | 9.56 | 10.98 | 9.64 |
| PRECT p99.99 | 110.21 mm/day | 30.79 | 32.35 | 31.79 | 33.05 | 31.29 |
| PRECT Rx (annual max) | 75.31 mm/day | 12.86 | 14.17 | 11.76 | 14.70 | 12.34 |
| PRECT freq > 8 mm/day | 0.0985 | 0.0133 | 0.0136 | 0.0142 | 0.0156 | 0.0132 |
Global-mean tail bias, model minus truth (closest to zero in bold):
| Tail statistic | Truth global mean | ckpt 69 | ckpt 70 | ckpt 52 | ckpt 36 | ckpt 46 |
|---|---|---|---|---|---|---|
| TREFHT p1 (cold) | 278.074 K | +0.076 | +0.255 | +0.324 | +0.326 | +0.387 |
| TREFHT p0.1 (cold) | 275.932 K | +0.116 | +0.375 | +0.410 | +0.389 | +0.440 |
| TREFHT p99 (hot) | 296.069 K | -0.120 | -0.225 | -0.112 | -0.166 | -0.208 |
| TREFHT p99.9 (hot) | 297.535 K | -0.027 | -0.159 | -0.002 | -0.092 | -0.094 |
| TREFHT TXx | 297.140 K | +0.042 | -0.075 | +0.117 | -0.0003 | +0.008 |
| TREFHT TNn | 276.765 K | +0.030 | +0.232 | +0.253 | +0.298 | +0.343 |
| PRECT p99 | 31.78 mm/day | -1.54 | -1.76 | -1.16 | -1.89 | -1.18 |
| PRECT p99.9 | 65.16 mm/day | -5.05 | -6.30 | -3.78 | -6.56 | -4.10 |
| PRECT p99.99 | 110.21 mm/day | -10.53 | -15.30 | -7.25 | -14.89 | -8.88 |
| PRECT Rx (annual max) | 75.31 mm/day | -8.54 | -10.17 | -6.69 | -10.26 | -7.38 |
| PRECT freq > 8 mm/day | 0.0985 | +0.0021 | +0.0024 | +0.0037 | +0.0039 | +0.0029 |
The tail score is the sum over these 11 statistics of the bias-RMSE z-score (across the five finalists), and the final ranking blends it 50/50 with the monthly-mean score: 69 > 70 > 52 > 36 > 46.
Variables
| Name | Meaning |
|---|---|
TREFHT | Reference-height (2 m) air temperature, in kelvin. The standard "surface air temperature". |
PRECT | Total precipitation (convective + large-scale). Native units are metres of liquid-water equivalent per 6-hourly step; mm/day = native x 4000. |
GMT | Global-mean temperature β the cos(latitude)-weighted spatial mean of TREFHT. |
Tail / extremes metrics (computed per grid point over the full 6-hourly record, then aggregated)
| Name | Meaning |
|---|---|
p99, p99.9, p99.99 | The 99th / 99.9th / 99.99th percentile of the 6-hourly values at each grid point β the hot tail for TREFHT, the heavy-rain tail for PRECT. p99.99 of 6-hourly data is roughly a once-a-year event. |
p1, p0.1 | The 1st / 0.1st percentile β the cold tail of TREFHT. |
TXx | Mean of the annual maxima of TREFHT β the typical hottest moment of a year ("T-max-max"). |
TNn | Mean of the annual minima of TREFHT β the typical coldest moment of a year ("T-min-min"). |
Rx | Mean of the annual maximum 6-hourly precipitation β the typical wettest 6 hours of a year. |
freq>~10mm/d, freq > 8 mm/day | Fraction of 6-hourly steps exceeding 2 mm per step (8 mm/day; the ~10mm/d label in the figures is a rounded name for the same threshold) β how often it rains hard, as opposed to how hard. |
TXx, TNn, and Rx follow the standard ETCCDI climate-extremes convention.
Error measures
| Name | Meaning |
|---|---|
| RMSE | Root-mean-square error vs the training target. All spatial averages are cos(latitude)-weighted, so grid cells are weighted by true area. |
| Global-mean bias | Model global mean minus truth global mean. Can be small even when the model is locally wrong, because errors of opposite sign cancel. |
| bias-RMSE | RMSE of the map of tail bias. Unlike global-mean bias, local errors cannot cancel. Related by RMSEΒ² = biasΒ² + (centred RMSE)Β², i.e. bias-RMSE already contains the global offset. |
| Centred RMSE | The part of the error left after removing the global-mean offset β i.e. how wrong the spatial pattern is, independent of any uniform shift. |
| Pattern correlation | Area-weighted spatial correlation between the model and truth maps: does the model put the extremes in the right places? |
| Trend error | Model minus truth linear trend in global-mean TREFHT over the record, in K/decade. Near zero means the free-running rollout is not drifting. |
| z-score | A metric re-expressed as standard deviations from the mean across the evaluated checkpoints. Lets metrics with different units be summed into one score. It is relative: a good z-score means "better than the other checkpoints", not "good in absolute terms". |
CAMulator is trained with a global energy fixer (an up/down rescaling of the
temperature field that closes the column energy budget against the prescribed
TOA and surface fluxes). This constraint is active at inference and is
enabled in the shipped inference_config.yaml; running without it is a
train/inference mismatch.
Enabling it costs a little accuracy relative to running unconstrained β it makes global-mean precipitation ~2% wetter, because rescaling temperature changes saturation vapour pressure and so perturbs precipitation indirectly β but it is the constraint the model was trained under, it is stable over 35-year rollouts (no drift, no NaN), and it is what makes the energy budget physically meaningful. All numbers and figures on this card are produced with the fixer active.
willychap/camulator
βββ README.md # this model card
βββ inference_config.yaml # ready-to-run config (= camulator_config.yml)
βββ checkpoint.pt00069.pt # default model (epoch 69); other top epochs alongside
βββ forcing_data/
β βββ b.e21.CREDIT_climate_cyclic_1yr_f32coords.nc # cyclic (default)
β βββ b.e21.CREDIT_climate_branch_1980_2014.nc # progressive/transient
βββ initial_conditions/
β βββ init_camulator_condition_tensor_*.pth # 69 ICs (Jan 1 & Jul 1, 1980/1981-2014)
βββ normalization/
β βββ mean_*.nc, std_*.nc # z-score
β βββ *statics*.nc # statics + latitude weights
βββ metadata/
β βββ camulator_metadata.yaml # output variable units / long-names
βββ figs/ # model-card figures
(The inference toolbox actually reads its units from the in-repo copy
climate/camulator_metadata.yaml; the metadata/ copy here is for reference.)
download_assets.py pulls these into the toolbox's ./assets/ for you.
CAMulator was trained on a CAM6 / ERA5-scaled climate dataset (1980-2014). The full training archive is not hosted here; the inputs needed to run the model (forcing, initial conditions, normalization, statics) are.
If you would like access to our training Zarr datasets, please email wchapman [at] colorado.edu.
@article{chapman2025camulator,
title = {CAMulator: Fast Emulation of the Community Atmosphere Model},
author = {Chapman, William E. and Schreck, John S. and Sha, Yingkai and
Gagne II, David John and Kimpara, Dhamma and Zanna, Laure and
Mayer, Kirsten J. and Berner, Judith},
journal = {arXiv preprint arXiv:2504.06007},
year = {2025},
doi = {10.48550/arXiv.2504.06007},
url = {https://arxiv.org/abs/2504.06007}
}