Downloads · 30 days
0
isharane/Credit-Risk-Loan-Default-Prediction
Credit-Risk-Loan-Default-Prediction is a machine learning model from isharane. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This project aims to predict the probability of credit default using a combination of Logistic Regression, Random Forest, and XGBoost classifiers combined in a soft voting ensemble. The dataset contains financial, dem…
Downloads · 30 days
0
Access
Public
Updated Aug 9, 2025
Repo size
53.8 MB
Likes
0
Public
Click a slice to open those files.
.pkl53.8 MB · 100%
From the Hugging Face model README
This project aims to predict the probability of credit default using a combination of Logistic Regression, Random Forest, and XGBoost classifiers combined in a soft voting ensemble. The dataset contains financial, demographic, and credit history attributes of loan applicants. The objective is to identify high-risk applicants and support better lending decisions.
Source: Give Me Some Credit.csv
Target variable: SeriousDlqin2yrs
0 → No serious delinquency in the past 2 years1 → Serious delinquency occurredKey Features:
RevolvingUtilizationOfUnsecuredLines – Ratio of credit card balance to credit limitNumberOfTime30-59DaysPastDueNotWorse – Count of 30-59 days late paymentsage – Age of the applicantNumberOfTimes90DaysLate – Count of 90+ days late paymentsDebtRatio – Monthly debt payments to income ratioMonthlyIncome – Monthly income of the applicantNumberOfOpenCreditLinesAndLoans – Number of open credit lines and loansData Preprocessing
Handling Missing Values: Median imputation for numerical columns.
Outlier Treatment: Clipping based on IQR method.
Feature Engineering:
DebtToIncomeRatio = RevolvingUtilizationOfUnsecuredLines / MonthlyIncome<30, 30-40, 40-50, 50-60, 60+)Scaling: StandardScaler applied to numerical features.
Class Imbalance Handling: Stratified train-test split to maintain target distribution.
Model Training
Three base models were trained and tuned using GridSearchCV with ROC-AUC as the scoring metric:
These were combined into a VotingClassifier with voting="soft" to leverage predicted probabilities from all models.
Model Performance
| Metric | Class 0 | Class 1 |
|---|---|---|
| Precision | 0.94 | 0.60 |
| Recall | 0.99 | 0.18 |
| F1-Score | 0.97 | 0.28 |
| Accuracy | 0.94 | |
| ROC-AUC Score | 0.864 |
Macro Avg F1: 0.62 Weighted Avg F1: 0.92
Feature Importance (SHAP Analysis)
The top features influencing the model’s predictions are: