Downloads · 30 days
0
sirunchained/diabet-analysis
diabet-analysis is a machine learning model from sirunchained. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This project develops machine learning models to predict diabetes onset using the Pima Indians Diabetes Database. The goal is to accurately classify patients based on diagnostic measurements while handling missing dat…
Downloads · 30 days
0
Access
Public
Updated Jun 20, 2026
Repo size
463 KB
Likes
0
Public
Click a slice to open those files.
.joblib130 KB · 96%
From the Hugging Face model README
This project develops machine learning models to predict diabetes onset using the Pima Indians Diabetes Database. The goal is to accurately classify patients based on diagnostic measurements while handling missing data, class imbalance, and outliers.
Dataset: 768 samples with 8 features + target (Outcome).
Task: Binary classification (diabetes: 1, no diabetes: 0).
Metrics: Accuracy and F1-score (as requested by the team).
The Support Vector Classifier (SVC) with default parameters (after scaling and SMOTE) achieved the highest F1-score with competitive accuracy:
| Metric | Score |
|---|---|
| F1-score | 65.55% |
| Accuracy | 73.38% |
But SVC currently dose not work with pipeline, I chosed KNN (with GridSearch) which has best scores after SVC:
| Metric | Score |
|---|---|
| F1-score | 65.57% |
| Accuracy | 72.73% |
| Model | F1‑Score | Accuracy |
|---|---|---|
| SVC (default) | 65.55% | 73.38% |
| KNN (GridSearch) | 65.57% | 72.73% |
| Random Forest (GridSearch) | 63.64% | 74.03% |
| Logistic Regression (GridSearch) | 63.16% | 72.73% |
| Decision Tree | 60.61% | 66.23% |
Note: GridSearchCV was used for hyperparameter tuning; SVC with default parameters (
C=1,kernel='rbf',gamma='scale') outperformed its tuned version (which had lower F1).
Handling Missing Values
0 as missing: Glucose, BloodPressure, SkinThickness, BMI.Glucose, BloodPressure, BMI.SkinThickness (due to high missing rate ~42%).Outlier Treatment
Class Imbalance
Feature Engineering
Glucose_BMI = Glucose × BMI, though it did not improve performance.Glucose and BMI were the most influential features.Age and Pregnancies (0.54).Insulin had ~48% missing values, SkinThickness ~42%.Try the application directly:
👉 Diabetes Prediction Demo
├── main.ipynb # Full pipeline: EDA, preprocessing, modeling, evaluation
├── diabetes.csv # Dataset
├── requirements.txt # Python dependencies
└── README.md # This file
Key libraries: pandas, numpy, matplotlib, seaborn, scikit-learn, imbalanced-learn, joblib.
jupyter notebook main.ipynb to explore the full analysis.SirUnchained
This project was developed as a solution for a diabetes prediction task, demonstrating a complete machine learning workflow.