Downloads · 30 days
0
DarwinDanish/exoplanet-classifier-stacking
exoplanet-classifier-stacking is a tabular classification model from DarwinDanish. Use it for the tabular classification task on the model card, and read the license before you ship it in a product. It is set up for sklearn. The card lists the license as mit.
This is a robust machine learning pipeline designed to classify Kepler Objects of Interest (KOIs). It determines whether a detected signal represents a real exoplanet or a false positive.
Downloads · 30 days
0
Access
Public
Updated Nov 23, 2025
Repo size
16.5 MB
Likes
0
Public
Click a slice to open those files.
.pkl16.5 MB · 100%
From the Hugging Face model README
This is a robust machine learning pipeline designed to classify Kepler Objects of Interest (KOIs). It determines whether a detected signal represents a real exoplanet or a false positive.
The model utilizes a Stacking Ensemble architecture, combining the predictions of three powerful gradient boosting frameworks (LightGBM, XGBoost, and CatBoost) and aggregating them using a final LightGBM meta-learner. It is specifically engineered to handle missing data (NaNs) in scientific datasets through a dual-strategy imputation pipeline.
This model is intended for astronomers, data scientists, and space enthusiasts who want to analyze Kepler mission data or similar photometric datasets. It predicts the "disposition" of a celestial object based on its physical properties.
To use this model, your input DataFrame must contain the following columns:
Critical Features:
koi_period: Orbital periodkoi_depth: Transit depthkoi_prad: Planetary radiuskoi_sma: Semi-major axiskoi_teq: Equilibrium temperaturekoi_insol: Insolation fluxkoi_model_snr: Signal-to-Noise RatioAuxiliary Features:
koi_time0bk, koi_duration, koi_incl, koi_srho, koi_srad, koi_smass, koi_steff, koi_slogg, koi_smetYou can load this model directly from the Hugging Face Hub using joblib and huggingface_hub.
pip install huggingface_hub joblib pandas scikit-learn lightgbm xgboost catboost
import joblib
import pandas as pd
from huggingface_hub import hf_hub_download
# 1. Download the model and label encoder
repo_id = "DarwinDanish/exoplanet-classifier-stacking"
model_path = hf_hub_download(repo_id=repo_id, filename="exo_stacking_pipeline.pkl")
encoder_path = hf_hub_download(repo_id=repo_id, filename="exo_label_encoder.pkl")
# 2. Load the artifacts
pipeline = joblib.load(model_path)
label_encoder = joblib.load(encoder_path)
# 3. Create sample data (Example: A likely planet candidate)
# Note: The model handles NaNs, so missing values are allowed.
data = {
'koi_period': [365.25],
'koi_depth': [1000.5],
'koi_prad': [1.02], # Earth radii
'koi_sma': [1.0], # AU
'koi_teq': [255.0], # Kelvin
'koi_insol': [1.0],
'koi_model_snr': [35.5],
# Aux features (can be mostly defaults or NaNs)
'koi_time0bk': [135.0],
'koi_duration': [4.5],
'koi_incl': [89.9],
'koi_srho': [1.0],
'koi_srad': [1.0],
'koi_smass': [1.0],
'koi_steff': [5700],
'koi_slogg': [4.5],
'koi_smet': [0.0]
}
df_new = pd.DataFrame(data)
# 4. Predict
prediction_index = pipeline.predict(df_new)
prediction_label = label_encoder.inverse_transform(prediction_index)
probabilities = pipeline.predict_proba(df_new)
print(f"Prediction: {prediction_label[0]}")
print(f"Confidence: {max(probabilities[0]):.4f}")
The model was trained using a robust preprocessing pipeline followed by a Stacking Classifier.
The pipeline splits features into two groups with different imputation strategies:
-999. Scaled via StandardScaler.median. Scaled via StandardScaler.Based on the base learners, the most critical features for classification were identified as:
koi_model_snr (Signal-to-Noise Ratio)koi_prad (Planetary Radius)koi_depth (Transit Depth)koi_period (Orbital Period)The model was evaluated on a held-out test set (20% of data) using stratified splitting.