Downloads · 30 days
0
Ajiboye/ny2017hospital_predict_model
ny2017hospital_predict_model is a tabular regression model from Ajiboye. Use it for the tabular regression task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This XGBoost regression pipeline predicts hospital Length of Stay (LOS) in days for inpatient admissions across New York State hospitals. The model was trained on 2.3+ million de-identified hospital discharge records…
Downloads · 30 days
0
Access
Public
Updated Feb 2, 2026
Repo size
4.8 MB
Likes
0
Public
Click a slice to open those files.
.pkl2.3 MB · 99%
From the Hugging Face model README
This XGBoost regression pipeline predicts hospital Length of Stay (LOS) in days for inpatient admissions across New York State hospitals. The model was trained on 2.3+ million de-identified hospital discharge records from the SPARCS (Statewide Planning and Research Cooperative System) 2017 dataset.
Intended Use: Support discharge planning, resource allocation, and patient expectation management by providing evidence-based LOS predictions with 95% confidence intervals.
✅ Clinical Decision Support
✅ Healthcare Operations
✅ Research & Analytics
❌ NOT for:
Input (13 features)
↓
┌─────────────────────────────────────────┐
│ HospitalDataCleaner │
│ - MDC description → code mapping │
│ - Target encoding (LOS_per_MDC) │
│ - Target encoding (LOS_per_severity) │
│ - One-hot encoding (categorical vars) │
│ - Feature alignment (312 columns) │
└─────────────────┬───────────────────────┘
↓
Encoded Features (312)
↓
┌─────────────────────────────────────────┐
│ XGBoost Regressor │
│ - n_estimators: 100 │
│ - max_depth: 6 │
│ - learning_rate: 0.1 │
│ - objective: reg:squarederror │
└─────────────────┬───────────────────────┘
↓
Predicted LOS (days)
Target Encoding:
LOS_per_MDC: Median LOS grouped by Major Diagnostic CategoryLOS_per_severity: Median LOS grouped by severity levelOne-Hot Encoding applied to:
Total Features After Encoding: 312
Source: Hospital Inpatient Discharges (SPARCS De-Identified) 2017
Cleaning Steps:
U)120+ to numeric value 120Data Split:
Length of Stay Statistics (days):
- Mean: 5.2
- Median: 3.0
- Std Dev: 6.8
- Min: 1
- Max: 120
- 25th percentile: 2
- 75th percentile: 6
| Metric | Training | Validation | Test |
|---|---|---|---|
| RMSE | X.XX days | X.XX days | X.XX days |
| MAE | X.XX days | X.XX days | X.XX days |
| R² | 0.XX | 0.XX | 0.XX |
| MAPE | X.X% | X.X% | X.X% |
Note: Update with your actual evaluation results
By Severity Level:
| Severity | MAE | Sample Size |
|---|---|---|
| 1 (Minor) | X.X days | ~800K |
| 2 (Moderate) | X.X days | ~900K |
| 3 (Major) | X.X days | ~500K |
| 4 (Extreme) | X.X days | ~150K |
By Diagnosis Group (Top 5):
| MDC Description | MAE | Sample Size |
|---|---|---|
| Circulatory System | X.X | ~300K |
| Respiratory System | X.X | ~250K |
| Digestive System | X.X | ~220K |
| Nervous System | X.X | ~180K |
| Pregnancy/Childbirth | X.X | ~200K |
Concordance with Expert Judgment:
pip install xgboost scikit-learn pandas numpy joblib
import joblib
import pandas as pd
# Load the full pipeline
pipeline = joblib.load('xgb_hospital_full_pipeline.pkl')
# Or load model + preprocessor separately
model = joblib.load('xgb_modelv1.pkl')
preprocessor = joblib.load('hospital_data_cleanerv1.pkl')
import pandas as pd
# Prepare input data (13 features)
patient_data = pd.DataFrame([{
'Hospital County': 'Kings',
'Facility Name': 'Mount Sinai Hospital',
'Age Group': '50 to 69',
'Gender': 'M',
'Race': 'White',
'Ethnicity': 'Not Span/Hispanic',
'Type of Admission': 'Emergency',
'Patient Disposition': 'Home or Self Care',
'APR MDC Code': 5, # Circulatory system
'APR MDC Description': 'Diseases and Disorders of the Circulatory System',
'APR Severity of Illness Code': 3,
'APR Medical Surgical Description': 'Medical',
'Payment Typology 1': 'Medicare',
'Emergency Department Indicator': 'Y'
}])
# Predict
predicted_los = pipeline.predict(patient_data)
print(f"Predicted LOS: {predicted_los[0]:.2f} days")
# Output: Predicted LOS: 4.47 days
# 1. Preprocess
X_processed = preprocessor.transform(patient_data)
# 2. Predict
predicted_los = model.predict(X_processed)
# 3. Calculate confidence interval (95%)
std_error = predicted_los[0] * 0.15
confidence_low = max(1.0, predicted_los[0] - 1.96 * std_error)
confidence_high = predicted_los[0] + 1.96 * std_error
print(f"Prediction: {predicted_los[0]:.1f} days")
print(f"95% CI: [{confidence_low:.1f}, {confidence_high:.1f}] days")
# Load multiple patients
patients_df = pd.read_csv('patient_admissions.csv')
# Predict for all
predictions = pipeline.predict(patients_df)
# Add to dataframe
patients_df['predicted_los'] = predictions
patients_df.to_csv('predictions_output.csv', index=False)
import matplotlib.pyplot as plt
# Get feature names from pipeline
feature_names = pipeline.named_steps['preprocessor'].get_feature_names_out()
# Get importance scores
importance = model.feature_importances_
# Sort and plot top 20
indices = importance.argsort()[-20:][::-1]
plt.figure(figsize=(10, 6))
plt.barh(range(20), importance[indices])
plt.yticks(range(20), [feature_names[i] for i in indices])
plt.xlabel('Feature Importance')
plt.title('Top 20 Most Important Features for LOS Prediction')
plt.tight_layout()
plt.show()
⚠️ Data Limitations:
⚠️ Model Limitations:
🔴 Demographic Biases:
🔴 Geographic Biases:
🔴 Clinical Biases:
✅ Implemented:
⚠️ Recommended for Production:
Clinical Context Required
Transparency with Patients
Avoid Discriminatory Use
Data Privacy
Model Governance
Demographic Parity (should be analyzed):
Example Analysis:
# Check prediction distributions by race
results_by_race = df.groupby('Race')['predicted_los'].describe()
print(results_by_race)
# Flag if mean predictions differ by >20% across groups
# (May indicate bias OR clinical differences - requires clinical review)
If you use this model in your research or application, please cite:
@misc{hospital_los_xgboost_2026,
author = {Ajiboye Toluwalase},
title = {Hospital Length of Stay Predictor - XGBoost Pipeline},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Ajiboye/hospital_predict_model}},
note = {Trained on SPARCS NY 2017 dataset}
}
Data Source Citation:
New York State Department of Health. (2017). Hospital Inpatient Discharges
(SPARCS De-Identified): 2017. https://health.data.ny.gov/
This repository contains:
hospital-los-xgboost/
├── xgb_hospital_full_pipeline.pkl # Complete pipeline (recommended)
├── xgb_modelv1.pkl # XGBoost model only
├── hospital_data_cleanerv1.pkl # Preprocessor only
├── feature_names.pkl # Expected 312 feature names
├── README.md # This model card
├── requirements.txt # Python dependencies
Total size: ~15 MB (compressed)
This model is released under the MIT License.
MIT License
Copyright (c) 2025 [Ajiboye Toluwalase]
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT.
⚕️ Remember: This model is a tool to support healthcare professionals, not replace them. Always involve clinical expertise in patient care decisions.
Last updated: February 2026