Downloads · 30 days
0
reefzehavi/readme2
readme2 is a machine learning model from reefzehavi. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This project applies end-to-end Machine Learning techniques to predict whether an American company will go bankrupt based on financial data. The dataset contains financial attributes (X1-X18) for thousands of companie…
Downloads · 30 days
0
Access
Public
Updated Dec 11, 2025
Repo size
39.6 MB
Likes
0
Public
Click a slice to open those files.
.pkl27.7 MB · 68%
From the Hugging Face model README
This project applies end-to-end Machine Learning techniques to predict whether an American company will go bankrupt based on financial data. The dataset contains financial attributes (X1-X18) for thousands of companies. The project follows a structured Data Science pipeline:
X1).Watch the full project walkthrough and code explanation here: (https://www.loom.com/share/c0e6ce867ce0412cba909c3406707ff3)
We started by analyzing the raw data structure, distributions, and correlations.
We analyzed the correlation between different financial features. Strong correlations helped us identify redundant features and potential interactions.

The original dataset labels ('alive' vs 'failed') were highly imbalanced.

To improve model performance, we engineered new features:
X16 * X7).X2 / X3).Cluster_ID as a new feature for the supervised models.
Goal: Predict the continuous variable X1 (Net Income) using the engineered features.
We trained and compared three models: Linear Regression (Ridge), Random Forest Regressor, and Gradient Boosting.
| Model | RMSE (Lower is better) | R² Score (Higher is better) |
|---|---|---|
| Ridge Regression 🏆 | 0.012 | 0.9999 |
| Random Forest | 888.67 | 0.934 |
| Gradient Boosting | 906.24 | 0.931 |
The Ridge Regression model achieved a near-perfect R2 score, indicating a strong linear relationship in the financial features.
Below is the Feature Importance showing the most influential predictors:

Goal: We converted the problem into a classification task by splitting the target X1 by its Median.
Using the median split ensured a perfectly balanced dataset for training:

We compared Logistic Regression, XGBoost, and Random Forest.
| Model | F1-Score | Accuracy | Recall (Class 1) |
|---|---|---|---|
| Random Forest 🏆 | 0.9976 | 0.9976 | 0.997 |
| XGBoost | 0.9970 | 0.9970 | 0.996 |
| Logistic Regression | 0.9931 | 0.9931 | 0.992 |
The Random Forest model outperformed others.
Confusion Matrix: The model made very few errors on the test set.

This repository contains all necessary files to reproduce the results:
intro_to_data_science_2.ipynb: The complete Python notebook with code, analysis, and visualizations.american_bankruptcy.csv: The dataset used for training and testing.best_regression_model.pkl: The trained Ridge Regression model.best_classification_model.pkl: The trained Random Forest Classification model.*.png: Visualization images generated during the process.Submitted as part of the Data Science Course Assignment.