Downloads · 30 days
20
31% of all-time downloads
rdinnager/tabpfn-sdm-finetuned
tabpfn-sdm-finetuned is a tabular classification model from rdinnager. Use it for the tabular classification task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as other.
Downloads · 30 days
20
31% of all-time downloads
All-time downloads
65
Public
Repo size
86.6 MB
Likes
1
Public
Click a slice to open those files.
.pt85.9 MB · 99%
From the Hugging Face model README
Built with PriorLabs-TabPFN
Finetuned TabPFN models for binary species distribution modeling (SDM). These models were adapted from the pretrained TabPFN foundation model to improve performance on presence/absence classification tasks common in ecology and conservation biology.
| File | Description | Epochs | Val ROC-AUC | Val PR-AUC |
|---|---|---|---|---|
tabpfn-sdm-nonspatial.pt | Standard (non-spatial) train/test evaluation | 100 | 0.747 | 0.261 |
tabpfn-sdm-spatial.pt | Spatially-separated train/test evaluation | 50 | 0.653 | 0.144 |
The non-spatial variant is trained and evaluated using standard random train/test splits. The spatial variant uses 10km buffer-based spatial separation between training and test data, which is more realistic for ecological applications but results in a harder prediction task.
Each model was finetuned in two steps:
disdat R package (standardized presence-only and presence-absence datasets)Prior-Labs/TabPFN-v2-clf)import torch
from tabpfn import TabPFNClassifier
from huggingface_hub import hf_hub_download
# Download model
model_path = hf_hub_download(
repo_id="rdinnager/tabpfn-sdm-finetuned",
filename="tabpfn-sdm-nonspatial.pt"
)
# Load finetuned model
device = "cuda" if torch.cuda.is_available() else "cpu"
clf = TabPFNClassifier(
ignore_pretraining_limits=True,
device=device,
n_estimators=8,
random_state=32639
)
clf._initialize_model_variables()
checkpoint = torch.load(model_path, map_location=device, weights_only=False)
clf.models_[0].load_state_dict(checkpoint["model_state_dict"])
clf.models_[0].eval()
# Fit on training data and predict
clf.fit(X_train, y_train)
probs = clf.predict_proba(X_test)[:, 1] # Probability of presence
The companion R code provides inference wrappers. See the TabPFN-SDM GitHub repository for the full pipeline.
source("R/run_tabpfn_finetuned.R")
result <- run_tabpfn_finetuned_ensemble(
train_dat = train_data,
test_dat = test_data,
model_path = "tabpfn-sdm-nonspatial.pt",
max_train_size = 1000L,
n_estimators = 8L
)
# result$test contains y_truth and y_pred columns
For best results, use ensemble inference which matches the training procedure:
This is implemented in run_tabpfn_finetuned_ensemble() (R) or can be replicated in Python using the utilities in the GitHub repo.
Each .pt file is a PyTorch checkpoint dictionary containing:
| Key | Description |
|---|---|
model_state_dict | Finetuned TabPFN model weights |
config | Training configuration (hyperparameters, data settings) |
history | Training history (loss, metrics per epoch) |
step1_path | Path to the Step 1 model used as initialization |
tabpfn Python package (v2) for inferenceIf you use these models, please cite the accompanying paper:
@article{dinnage2026niche,
title={A Niche in the Machine: The Promise of AI Foundation Models for Species Distribution Modeling},
author={Dinnage, Russell and Warren, Dan L.},
year={2026},
doi={10.32942/X2VQ10},
journal={EcoEvoRxiv},
url={https://ecoevorxiv.org/repository/view/11797/}
}
Please also cite the original TabPFN paper:
@article{hollmann2025tabpfn,
title={Accurate Predictions on Small Data with a Tabular Foundation Model},
author={Hollmann, Noah and M{\"u}ller, Samuel and Purucker, Lennart and
Krishnakumar, Arjun and K{\"o}rfer, Max and Hoo, Shi Bin and
Schirrmeister, Robin Tibor and Hutter, Frank},
journal={Nature},
year={2025},
publisher={Nature Publishing Group}
}
These finetuned model weights are distributed under the Prior Labs License v1.1, consistent with the base TabPFN model license. See the LICENSE file for full terms.