Downloads Β· 30 days
0
Danielhalali/diabetes-regression-model
diabetes-regression-model is a machine learning model from Danielhalali. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Downloads Β· 30 days
0
Access
Public
Updated May 5, 2026
Repo size
619 MB
Likes
0
Public
Click a slice to open those files.
.mp495 MB Β· 81%
From the Hugging Face model README
<video src="https://huggingface.co/Danielhalali/diabetes-regression-model/resolve/main/assignment 2 video.mp4" controls="controls" style="max-width: 720px;"></video>
This project was completed as part of a Data Science course assignment. The goal is to predict diabetes risk using lifestyle, demographic, and clinical data from nearly 100,000 individuals.
I explored the data, engineered new features, applied clustering, and trained multiple regression and classification models β comparing their performance and selecting the best one.
Research Question: Can we predict whether a person has diabetes based on their health and lifestyle data? And if so β which factors matter most?
diagnosed_diabetes (0 = Healthy, 1 = Diabetic)diabetes_risk_score converted to 3 balanced classes β Low Risk, Medium Risk, High RiskClass Distribution: The dataset is relatively balanced β 59% of participants are diagnosed with diabetes and 41% are not. <img src="./pic1.png" width="600"/>
Age Distribution: Most participants are between ages 30 and 70, with a peak around age 50. The distribution gives a good demographic overview of the dataset and shows we have representation across a wide age range. <img src="./pic2.png" width="600"/>
Feature Correlation Heatmap: The heatmap shows how strongly each pair of features is related. Dark red means a strong positive relationship β when one goes up, the other goes up too. Fasting glucose and HbA1c are the darkest red relative to the diabetes diagnosis column, meaning they are the strongest predictors. Most other features appear light-colored, indicating weak direct correlations. <img src="./pic3.png" width="600"/>
Does BMI predict diabetes? Yes β the boxplot clearly shows that diabetic patients have a significantly higher median BMI than healthy individuals. This confirms that obesity is a key risk factor for diabetes in this dataset.
<img src="./pic4.png" width="600"/>Does smoking affect diabetes? Partially β across all smoking groups (Never, Former, Current), there are more diabetic individuals than healthy ones. However, the proportion doesn't change dramatically between groups, suggesting smoking alone is not the primary driver of diabetes in this dataset.
<img src="./pic5.png" width="600"/>Does fasting glucose increase with age? Yes β the scatter plot with a regression line shows a clear upward trend from left to right. As age increases, fasting glucose levels tend to rise as well. This helps explain why diabetes risk increases with age.
<img src="./pic6.png" width="600"/>The baseline model was trained on the original features with default parameters.
| Metric | Value |
|---|---|
| MAE | 0.0832 |
| RMSE | 0.1366 |
| RΒ² | 0.9221 |
The baseline model explains 92% of the data. Feature importance shows that diabetes_stage and HbA1c are the strongest predictors β both well-known medical risk factors for diabetes.
Actual vs Predicted: The plot shows predictions clustered around the binary values (0 and 1). The red dashed line represents a perfect prediction β the closer the points are to the line, the more accurate the model is. <img src="./pic13.png" width="600"/>
Residuals Distribution: The residuals plot shows 3 peaks β this is expected since the target is binary (only 0 or 1). The largest peak is centered around zero, meaning most predictions were close to the actual values. <img src="./pic14.png" width="600"/>
Feature Importance: The chart shows which features influenced the model most. diabetes_stage dominates β which makes sense since it is directly related to the diagnosis. HbA1c and family history of diabetes follow as the next most important predictors. <img src="./pic15.png" width="600"/>
I created 5 new interaction features from existing columns:
| Feature | Description |
|---|---|
glucose_insulin_ratio | Fasting glucose divided by insulin β captures insulin resistance |
bmi_age_interaction | BMI multiplied by age β combined metabolic risk |
pulse_pressure | Systolic BP minus diastolic BP β cardiovascular health indicator |
cholesterol_hdl_ratio | Total cholesterol divided by HDL β classic cardiovascular risk index |
glycemic_load | HbA1c multiplied by fasting glucose β combined glycemic risk |
All numeric features were then normalized using StandardScaler so that features with large values don't dominate the model over features with small values.
I used the Elbow Method to determine the optimal number of clusters β the graph shows a clear "elbow" at k=4, meaning adding more clusters beyond that doesn't significantly improve the groupings. <img src="./pic7.png" width="600"/>
I then visualized the clusters using PCA, which reduces all 31 dimensions to just 2 so we can draw a scatter plot. Each dot is one person, and each color is a different cluster. <img src="./pic8.png" width="600"/>
The clusters revealed 4 distinct health profiles:
| Cluster | Age | BMI | Glucose | HbA1c | Diabetic % |
|---|---|---|---|---|---|
| 0 β Young & Healthy | 40 | 23.5 | 99 | 5.76 | 12% |
| 1 β High Risk / Diabetic | 45 | 24.3 | 116 | 7.01 | 96% |
| 2 β Older High Risk | 62 | 27.8 | 124 | 7.26 | 98% |
| 3 β Middle-aged Moderate Risk | 56 | 27.2 | 105 | 6.02 | 30% |
The cluster ID and distance to centroid were added as new features for the models.
After feature engineering and clustering, I retrained three models on the enriched dataset:
| Model | RΒ² | MAE | RMSE |
|---|---|---|---|
| Baseline Linear Regression | 0.9221 | 0.0832 | 0.1366 |
| Improved Linear Regression | 0.9337 | 0.0770 | 0.1260 |
| Random Forest | 0.9987 | 0.0006 | 0.0174 |
| Gradient Boosting β | 0.9988 | 0.0008 | 0.0166 |
π Regression Winner: Gradient Boosting (RΒ² = 0.9988)
Random Forest and Gradient Boosting dramatically outperform Linear Regression because they can capture non-linear relationships between features β something Linear Regression cannot do. The engineered features (especially glycemic_load and glucose_insulin_ratio) appear in the top 15 most important features, confirming that feature engineering improved the models.
I converted the continuous diabetes_risk_score into 3 balanced classes using Quantile Binning (pd.qcut), each containing approximately 33% of the data. <img src="./pic16.png" width="600"/>
I focused on Recall as the key metric β in a medical context, a False Negative (missing a high-risk patient) is far more dangerous than a False Positive (incorrectly flagging a healthy person).
| Model | Accuracy |
|---|---|
| Logistic Regression | 0.80 |
| Random Forest | 0.93 |
| Gradient Boosting β | 0.94 |
π Classification Winner: Gradient Boosting (Accuracy = 0.94)
The confusion matrices show that Logistic Regression struggles most with the Medium Risk class, often confusing it with Low or High Risk. Gradient Boosting has significantly fewer off-diagonal errors across all classes. <img src="./pic10.png" width="600"/> <img src="./pic11.png" width="600"/> <img src="./pic12.png" width="600"/>
An interactive scatter plot showing fasting glucose vs BMI colored by cluster, with point size representing HbA1c level. Hovering over each point reveals age and diabetes diagnosis. The plot clearly shows that higher glucose and BMI correspond to the high-risk clusters. <img src="./bonus1.png" width="600"/>
5-fold cross validation on a sample of 10,000 records:
All 5 folds scored between 0.928 and 0.934, confirming the model is consistent and not overfitted. <img src="./bonus2.png" width="600"/>
The learning curve shows that as training size increases, validation accuracy steadily improves from 0.88 to 0.93. The training and validation curves are converging β confirming the model generalizes well and is not overfitting. Adding more data would likely continue to improve performance. <img src="./bonus3.png" width="600"/>
Assignment_2.ipynb β Full Python notebook with all analysis, models, and visualizationsregression_model.pkl β Winning regression model (Gradient Boosting, RΒ² = 0.9988)classification_model.pkl β Winning classification model (Gradient Boosting, Accuracy = 0.94)Diabetes_and_LifeStyle_Dataset .csv β Dataset with manually introduced missing values for data cleaning practice