Downloads · 30 days
0
Daniel-1109/ProsperLoan_Models
ProsperLoan_Models is a machine learning model from Daniel-1109. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
<video controls width="100%" <source src="https://huggingface.co/Daniel-1109/ProsperLoanModels/resolve/main/Presentation-DanielTenenboim-ProsperLoan.mp4?download=true" type="video/mp4" Your browser does not support th…
Downloads · 30 days
0
Access
Public
Updated Dec 8, 2025
Repo size
175 MB
Likes
0
Public
Click a slice to open those files.
.csv86.5 MB · 64%
From the Hugging Face model README
https://colab.research.google.com/drive/10Rt5BJyX-tcX8mHaiFN_KVzcPWG23YYk#scrollTo=w84cR3AZIU0e
The dataset used in this project is the Prosper Loan Dataset, obtained from Kaggle. Prosper is a U.S. peer to peer lending platform where individuals can apply for and invest in personal loans. The dataset contains detailed borrower-level financial, demographic, credit, and loan-performance information collected throughout the loan application and servicing process.
Each record represents a single loan, described through a structured set of borrower characteristics, credit history indicators, and loan attributes.
Borrower & Demographic Attributes- Employment Status, Occupation, Stated Monthly Income, Employment Length.
Credit-Related Features- Credit Score Range, Number of Credit Lines, Delinquency History, Public Records, Revolving Credit Balance, Bankcard Utilization, and other creditworthiness indicators.
Loan Characteristics- Loan Amount, Loan Term, BorrowerAPR (effective interest rate), Monthly Payment, Percent Funded, Number of Investors, Prosper Rating, Origination Date.
Economic & Performance Indicators- Debt to Income Ratio, Collection Fees, Principal Payments, Loss Metrics recorded during the loan lifecycle.
The goal of this analysis is to identify which borrower, credit and loan attributes influence the effective interest rate (BorrowerAPR) in the Prosper lending platform.
By examining variations in income levels, credit scores, debt to income ratios, employment status and credit history indicators, we aim to uncover the key factors that shape risk based loan pricing.
The analysis focuses on how core financial and demographic attributes such as stated monthly income, credit score ranges, debt to income ratio, revolving credit balance, delinquency history and credit history length reflect borrower stability and creditworthiness. Understanding these relationships helps explain the underlying drivers of interest rate determination.
The target variable in this analysis is BorrowerAPR, the effective interest rate assigned to each loan and a direct reflection of borrower risk. The goal is to understand which financial and credit factors drive APR higher or lower.
Later, this continuous variable is also grouped into ranges for classification, allowing us to explore how borrower characteristics map to different APR levels.
The Prosper dataset underwent a comprehensive cleaning process to ensure high data quality and prevent leakage in later modeling. Columns with more than 70 percent missing values were removed because they provided limited usable information and could not be reliably imputed. This included outdated or sparsely populated credit grade fields. Additional variables such as ClosedDate and Prosper’s internal model outputs were dropped because they reveal information from after the APR decision and would introduce leakage into the analysis.
Columns with moderate missingness, including Prosper rating fields and estimated performance metrics, were retained for exploration since their relevance had not yet been fully evaluated. Missing categorical values were filled with an explicit “Unknown” label, while numerical fields were imputed using the median to preserve distributional integrity in the presence of skewed financial data. Features with very small fractions of missing values followed the same strategy, ensuring completeness without distorting statistical patterns.
We confirmed that no duplicate records existed, and key categorical fields were standardized by trimming spaces and normalizing text formats. Sanity checks were performed to detect unrealistic values. Structurally impossible entries such as percent funded values above one and excessively large employment duration values were removed, while extreme but plausible financial values such as high revolving credit balances were preserved for later outlier analysis.
The initial sanity checks guided which variables required focused outlier inspection, leading us to examine stated monthly income, debt to income ratio and revolving credit balance. To improve readability, some outlier plots use a capped or broken x axis. This helps display the main distribution clearly while still showing extreme values, without altering the underlying data.
<img src="https://cdn-uploads.huggingface.co/production/uploads/6904aebe18a1ba17d9435d1e/Iaeb_RUApV3XrE7DjkIx1.png" width="1000"> For stated monthly income, most extreme values were legitimate high earners and were retained. A sanity threshold of 250,000 was applied to flag implausible values, and only 6 records were removed. <img src="https://cdn-uploads.huggingface.co/production/uploads/6904aebe18a1ba17d9435d1e/MT7qOoPD9phrYJHtJWVBX.png" width="1000"> For debt to income ratio, statistical outliers began around 0.8, but values up to 3 are still economically reasonable, so they were kept. Only ratios above 3 were considered unrealistic and removed. <img src="https://cdn-uploads.huggingface.co/production/uploads/6904aebe18a1ba17d9435d1e/R3PQFz3Hv-ZaafLsYbHmz.png" width="1000"> For revolving credit balance, most borrowers fell within a normal range, but a threshold of 300,000 identified values unlikely to reflect true consumer credit behavior. Less than 0.2 % of records exceeded this limit and were removed.These steps allowed us to preserve meaningful financial variation while eliminating values that were clearly implausible or indicative of data errors.
Across the three features examined for outliers (StatedMonthlyIncome, DebtToIncomeRatio and RevolvingCreditBalance), the IQR method flagged many observations as statistical outliers, though most represented legitimate financial variation. Instead of removing all flagged values, we applied domain based sanity thresholds to distinguish plausible extremes from economically impossible entries. Only 610 records (less than 0.61 % of the dataset) exceeded these limits and were removed. This selective approach preserves meaningful high value cases while improving overall data quality.
• Borrowers earn 4600-6800 per month and have 5-6 years of employment, indicating generally stable income.
• Median revolving payments (271) and balances (8500) show typical consumer credit use.
• Bankcard utilization averages 60%, reflecting moderate reliance on available credit.
• Borrowers hold 9–10 active credit lines, about 25 total lines, and 6 revolving accounts, showing consistent engagement with credit products.
• BorrowerAPR is usually 21%, interest rates 18–19%, and most loans are 36 months, with purposes concentrated in debt consolidation, home improvement and business.
• Most borrowers have 0–1 recent inquiries, 4–7 total inquiries, and rare delinquencies, indicating clean credit histories.
• Borrowers with stable employment receive lower APRs, while those with less reliable income face higher rates due to increased perceived risk.
To keep the README concise, only a selected subset of visualizations is included here. The full analysis contains additional plots and deeper exploratory visuals, all of which are available in the accompanying Colab notebook.
The distribution of BorrowerAPR shows that most loans fall within the 15%–25% range, with a smaller right-skewed tail representing higher risk borrowers who receive significantly higher rates. This pattern is typical in consumer lending and reflects how lenders price loans based on perceived borrower risk.
This chart displays the most common loan purposes among borrowers, using Prosper’s official category definitions. Debt Consolidation dominates the platform, accounting for over half of all loans, followed by categories such as Other, Not Available, Home Improvement and Business. The distribution highlights the primary financial motivations driving borrowers to seek funding.
The heatmap reveals strong correlations among BorrowerAPR, BorrowerRate and LenderYield, which represent different forms of the interest rate. LoanOriginalAmount is closely tied to MonthlyLoanPayment, as larger loans require larger payments.
Investor count shows moderate positive correlation with loan size and a negative relationship with APR, indicating that larger, lower risk loans attract more investors. Overall, the heatmap illustrates the structural relationships that shape loan pricing and borrower risk.
Only a subset of research visualizations is shown here for readability. Full analysis with all six research questions and plots is available in the Colab notebook.
Across all analyses, lenders appear to price APR using multiple signals of borrower risk. Lower income, smaller loans, higher DTI, and prior delinquencies all correspond to higher APRs. Credit score, approximated from each borrower's reported range, shows a strong inverse relationship with APR. Together, these factors form a consistent risk framework in which APR rises as financial stability decreases.
The goal of this regression task is to predict a borrower's APR based on financial, behavioral, and credit-related features available at the time the loan was issued.
For the baseline regression model, we selected only numerical features available at the time of loan origination to create a simple, unbiased starting point. Limiting the model to numeric variables avoids encoding steps and reflects the raw predictive signal in the data. We excluded identifiers, dates, categorical fields, and any post loan performance information to prevent leakage and ensure that all inputs represent information truly known when the APR was assigned.
We defined BorrowerAPR as the target variable (y) and the selected numerical predictors as the feature matrix (X). The data was then split into training and testing sets, reserving 20% for unbiased evaluation while ensuring reproducibility with a fixed random seed.
A simple Linear Regression model was trained on the training set to form a baseline, allowing us to assess the inherent predictive signal in the numerical features before applying more advanced modeling techniques.
We evaluated the baseline Linear Regression model using MAE, MSE, RMSE, and R². RMSE was used as the primary metric because it penalizes large errors and is expressed in the same units as APR, making it intuitive to interpret. To contextualize model performance, we compared the RMSE (0.0615) with the actual distribution of BorrowerAPR values. Given that most APRs fall between 0.15 and 0.28, this error level is reasonable for an unengineered baseline and highlights clear room for improvement through feature engineering and more advanced models.
MAE: 0.0487 | MSE: 0.0038 | RMSE: 0.0615 | R²: 0.4115
Although the model is simple and cannot capture non linear patterns, its coefficient directions align well with financial logic. This provides a clear baseline understanding of which borrower characteristics most strongly influence APR before moving on to more advanced models.
We removed additional columns that cause data leakage because they are generated after APR is determined or depend on Prosper’s internal scoring model. We also removed identifier columns that provide no predictive value. These features were not used in the baseline model, and excluding them ensures a cleaner and leakage free dataset for training the advanced models.
To strengthen the model, we engineered several features that capture combined financial behaviors and relationships not visible from individual variables alone.
• CreditHistoryYears - measures the length of each borrower's credit history at the time of loan origination.
• CreditScore - midpoint of the reported lower and upper credit score bounds, creating a single interpretable score.
• IncomeStabilityRatio - compares monthly income to debt burden, indicating repayment capacity.
• RevolvingBalanceBurden - reflects how heavy revolving balances are relative to income, signaling financial strain.
• DelinquencyPressureIndex - combines long term and current delinquencies into a stronger risk indicator.
• LoanToIncomeRatio - quantifies how large the requested loan is relative to income, capturing repayment pressure.
This compact feature set provides the model with richer risk-related signals while remaining fully interpretable.
To prepare categorical data for modeling, we converted all non-numeric fields into numerical representations. Boolean variables were mapped to 0/1, meaningful date components were extracted from loan timing fields, and one hot encoding was applied to categories such as employment status and loan quarter.
For high cardinality features, we used frequency encoding to avoid creating excessive columns. These transformations preserve the information in each variable while producing a clean, model ready numerical dataset.
The numerical features in the dataset span very different scales (for example, income values in the thousands vs. ratios between 0 and 1). To ensure that all variables contribute equally during training, we apply StandardScaler, which standardizes each feature to have mean 0 and standard deviation 1. This prevents large scale features from dominating the model and helps regression algorithms converge more effectively.
In this section, we use K-Means clustering to group borrowers into behavioral risk segments.
The clustering is based on standardized financial features that capture credit usage, delinquency history, and overall credit activity, so that each cluster reflects a distinct borrower risk profile.
Based on the Elbow and Silhouette results, we fit K-Means with k = 3.
| Feature | Cluster 0 | Cluster 1 | Cluster 2 |
|---|---|---|---|
| CreditScore | 0.078 | -0.805 | 0.227 |
| DebtToIncomeRatio | -0.072 | -0.405 | 0.355 |
| DelinquenciesLast7Years | -0.272 | 1.754 | -0.274 |
| RevolvingCreditBalance | -0.154 | -0.431 | 0.545 |
| TradesNeverDelinquent (%) | 0.17 | -1.61 | 0.41 |
To understand what differentiates the clusters, we examine the mean standardized feature values for each group. Cluster 0 shows average and stable financial behavior, Cluster 1 represents borrowers with the weakest credit indicators and highest delinquency levels, and Cluster 2 reflects higher income and activity but low delinquencies. These numerical patterns confirm the PCA separation and validate that the clusters represent meaningful borrower profiles.
In addition, Each borrower now gets three numbers measuring how close they are to each cluster's centroid. Smaller distance means more similar behavior. These continuous signals give the model richer information than a single cluster label.
We applied scaling again after creating the new cluster based features because their numeric ranges differ from the existing variables.
Re-scaling ensures that all features operate on a comparable scale and prevents the model from unintentionally overweighting the newly added features.
We exclude RiskCluster from scaling because it is a categorical group label (0/1/2), treating it as a continuous value would distort its meaning and mislead the model. Rescaling at the final stage of feature engineering is standard and ensures a clean, model ready dataset.
Linear Regression shows the same performance as the baseline because scaling and feature engineering do not add new linear signal for it to learn. Since the model can only capture linear relationships, its predictions remain unchanged, indicating that more advanced models are needed to benefit from the engineered features.
Ridge Regression significantly improves over Linear Regression: all errors decrease (MAE, MSE, RMSE) and R² rises from 0.42 to 0.53. By shrinking coefficients and handling multicollinearity, Ridge makes better use of the engineered and cluster-based features, resulting in a more stable and accurate model.
MAE: 0.0436 | MSE: 0.0030 | RMSE: 0.0549 | R²: 0.5311
Ridge Regression delivers a clear improvement over the baseline model by reducing errors and providing more stable, reliable predictions. Its regularization helps handle multicollinearity, allowing the model to better leverage the engineered features.
The XGBoost model delivers a substantial performance boost, achieving much lower errors and a high R² of 0.75. This indicates that XGBoost captures non-linear relationships and complex interactions that linear models cannot. Overall, it provides the strongest and most accurate APR predictions among all tested models.
MAE: 0.0300 | MSE: 0.0016 | RMSE: 0.0401 | R²: 0.7502
The XGBoost residual plot shows residuals tightly centered around zero with no strong pattern, indicating low bias and more stable errors across all APR levels. Compared to Ridge, the systematic under and over prediction patterns largely disappear, reflecting XGBoost's ability to capture nonlinear relationships and produce more reliable predictions.
XGBoost uncovers a distinct pattern of feature importance, highlighting relationships that linear models cannot capture.
The model ranks CreditScore as the most influential factor, followed closely by several LoanYear features, suggesting that historical market conditions and lending environments significantly shape APR outcomes.
Features like AvailableBankcardCredit reflect borrowers' available credit and financial stability, further contributing meaningful predictive value.
Overall, XGBoost identifies nonlinear interactions and temporal effects that traditional linear models tend to overlook.
| Model | MAE | MSE | RMSE | R² |
|---|---|---|---|---|
| Linear Regression | 0.0485 | 0.0037 | 0.0615 | 0.418 |
| Ridge Regression | 0.0436 | 0.0030 | 0.0549 | 0.531 |
| XGBoost Regression | 0.0300 | 0.0016 | 0.0401 | 0.750 |
XGBoost delivers by far the best performance among all models tested. Unlike Linear Regression and Ridge, which captured only simple linear trends, XGBoost learned nonlinear relationships and feature interactions, achieving much higher accuracy (R² = 0.75) and significantly lower prediction errors.
Its strong results were powered by the earlier feature engineering: credit behavior features, temporal loan indicators, and cluster based risk variables became highly influential, enabling the model to detect complex borrower risk patterns.
Overall, the combination of rich engineered features and a nonlinear model makes XGBoost the most effective and reliable predictor of BorrowerAPR.
We converted APR into three classes using quantile binning to create a balanced distribution across low, medium, and high APR levels. This preserves the natural ordering of APR while avoiding class imbalance, ensuring that classification models learn meaningful risk tiers.
We trained three classification models - Logistic Regression, Random Forest, and XGBoost and evaluated their performance using confusion matrices to understand how well they distinguish between the three APR classes (low, medium, high).
Across all three models, XGBoost delivers the most accurate and consistent predictions. Logistic Regression struggles with non linear patterns and often confuses neighboring APR classes. Random Forest improves class separation but still shows noticeable misclassification in boundary cases. XGBoost achieves the clearest diagonal pattern in the confusion matrix, indicating far fewer false negatives and false positives, and demonstrates the strongest ability to capture complex relationships learned during feature engineering.
Overall, XGBoost is the best performing classifier, offering the highest reliability for distinguishing APR risk tiers.
Higher APR segments appear strongly associated with instability indicators such as high DTI and delinquency pressure. Such patterns highlight the ethical consideration that financially stressed borrowers often pay disproportionately higher rates. Models predicting APR must therefore be used responsibly to avoid reinforcing disadvantage.
Through iterative training, I observed how significantly model performance improves when the features are meaningful. The linear models showed only limited gains, but after adding richer engineered features—temporal signals, credit behavior variables, and clustering, the more advanced models, especially XGBoost, improved dramatically. This reinforced the idea that strong feature engineering often matters more than the choice of algorithm.
I initially began the project with a different dataset and invested quite a lot of work exploring and preparing it, but eventually realized it wouldn’t support the type of analysis I wanted, so I decided to switch. Later in the process, I also learned an important pipeline lesson: after accidentally overwriting the train/test split following feature engineering, all model performance dropped, and it took hours to trace the issue. This reinforced how essential clean, consistent workflow steps are for reliable modeling.