Downloads · 30 days
0
zakaneki/wildfirez
wildfirez is a machine learning model from zakaneki. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Predicting wildfire size classes using machine learning on the FPA FOD (Fire Program Analysis Fire-Occurrence Database) containing 1.88 million US wildfire records from 1992-2015.
Downloads · 30 days
0
Access
Public
Updated Jan 4, 2026
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.sqlite796 MB · 66%
From the Hugging Face model README
Predicting wildfire size classes using machine learning on the FPA FOD (Fire Program Analysis Fire-Occurrence Database) containing 1.88 million US wildfire records from 1992-2015.
This project builds an ordinal classification model to predict fire size categories:
wildfires/
├── config/
│ ├── __init__.py # Package init
│ └── config.py # Configuration settings
├── data/
│ └── processed/ # Processed parquet files (train/test splits)
├── models/ # Saved model artifacts
│ ├── best_params.json # Tuned hyperparameters
│ ├── model_metadata.joblib # Feature names and metrics
│ └── wildfire_model.txt # Trained LightGBM model
├── reports/
│ └── figures/ # Visualizations and metrics
├── scripts/
│ ├── 01_extract_data.py # Extract SQLite → Parquet
│ ├── 02_eda.py # Exploratory data analysis
│ ├── 03_preprocess.py # Data preprocessing
│ ├── 04_feature_engineering.py # Feature creation
│ ├── 05_train_model.py # Model training
│ ├── 06_evaluate.py # Model evaluation
│ └── 07_predict.py # Prediction pipeline
├── run_pipeline.py # Run full or partial pipeline
├── requirements.txt # Dependencies
├── .gitignore # Git ignore rules
└── README.md
FPA_FOD_20170508.sqlite)python -m venv venv
venv\Scripts\activate # Windows
# source venv/bin/activate # Linux/Mac
pip install -r requirements.txt
Using the pipeline runner (recommended):
# Run full pipeline
python run_pipeline.py
# Skip EDA step
python run_pipeline.py --skip-eda
# Run with hyperparameter tuning
python run_pipeline.py --tune
# Resume from a specific step (1-7)
python run_pipeline.py --from-step 5
Or execute scripts individually:
# 1. Extract data from SQLite
python scripts/01_extract_data.py
# 2. Exploratory data analysis (generates plots)
python scripts/02_eda.py
# 3. Preprocess data
python scripts/03_preprocess.py
# 4. Feature engineering
python scripts/04_feature_engineering.py
# 5. Train model (add --tune for hyperparameter tuning)
python scripts/05_train_model.py
# python scripts/05_train_model.py --tune # With Optuna tuning
# 6. Evaluate model
python scripts/06_evaluate.py
# 7. Make predictions
python scripts/07_predict.py --lat 34.05 --lon -118.24 --state CA --cause "Lightning"
For ordinal classification, we prioritize:
After running the pipeline:
data/processed/: Parquet files for train/test splitsmodels/wildfire_model.txt: Trained LightGBM modelmodels/model_metadata.joblib: Feature names and metricsreports/figures/: Visualizations (confusion matrix, SHAP plots, etc.)Fire Program Analysis Fire-Occurrence Database (FPA FOD)
This project uses publicly available government data.