Downloads · 30 days
0
MayaKitzis/smoking-prediction-model
smoking-prediction-model is a machine learning model from MayaKitzis. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
0
Access
Public
Updated Dec 9, 2025
Repo size
446 MB
Likes
0
Public
Click a slice to open those files.
.mp4369 MB · 82%
From the Hugging Face model README
Video Link:
Goal: Analyze factors related to youth smoking and drug use and build models that
Dataset Link: https://www.kaggle.com/datasets/waqi786/youth-smoking-and-drug-dataset
Smoking_Prevalence – continuous score indicating level of smoking.Exploratory analysis was performed to understand the distribution of key variables and explore relationships relevant to smoking and drug use.
Smoking_Prevalence is roughly centered around ~27–28 with moderate spread.Peer_Influence, Media_Influence, Mental_Health are on a 1–10 scale.Key exploratory plots included:
Smoking Prevalence vs. Peer Influence
Line/point plots of average smoking across peer-influence levels showed a weak but
mostly increasing trend: higher peer influence is loosely associated with higher smoking levels.
Drug Experimentation vs. Family Background
Boxplot suggested that higher family risk (worse family background) tends to be associated
with slightly higher levels of drug experimentation, but the relationship is noisy.

Peer Influence by Age Group A boxplot showed how peer-influence scores are distributed across different age groups. Although the median peer-influence levels remain fairly consistent across ages, younger groups (10–24) displayed slightly wider variability. Overall, peer influence does not show strong age-related trends but remains relatively stable across the population.
The initial regression model aimed to predict smoking prevalence from demographic, behavioral, and social factors.
Conclusion of the feature importance:
The baseline linear regression model performed poorly. The near-zero explanatory power suggests that a simple linear relationship does not adequately capture the complexity of smoking behavior. This motivated more advanced modeling approaches and feature engineering.
To improve model performance, several engineered features were created:
Scaling
Polynomial Feature
Peer_Influence_Sq = Peer_Influence^2Combined Feature
Family_Community_Support – sum/combination of family and community support variables.Clustering Feature
Cluster_ID.These engineered features were re-used both for regression and classification tasks.
The improved model revealed:
This model uncovered patterns the baseline model failed to capture.

Random Forest captured complex nonlinear interactions that differ from the linear model.
Common regression metrics were used:
| Model | MAE | MSE | RMSE | R² |
|---|---|---|---|---|
| Baseline Linear Regression | ~11.1 | ~167 | ~12.9 | ≈ 0 or slightly negative |
| Improved Linear Regression | ~11.1 | ~167 | ~12.9 | ≈ 0 |
| Random Forest Regressor | ~11.3 | ~176 | ~13.3 | ≈ 0 (slightly negative) |
| Gradient Boosting Regr. | ~11.1 | ~168 | ~12.98 | ≈ 0 (slightly negative) |
All models reached similar error levels, and the R² values near zero indicate that the dataset contains weak predictive signals for smoking prevalence.
The baseline linear regression model fails to capture meaningful patterns in the data, predicting nearly constant values regardless of the actual smoking prevalence
In the improved Linear Regression model, the most influential coefficients were:
Peer_Influence_Sq, Mental_Health, Media_Influence, Community_SupportFamily_Background (suggesting that a stronger family background tends to reduce smoking)In Random Forest and Gradient Boosting,
Drug_Experimentation was the most important predictor, followed by Social_Risk and Family_Community_Support.
Despite overall limited predictive strength, Gradient Boosting consistently achieved the lowest error metrics among the tested regression models. Therefore, Gradient Boosting is selected as the winning regression model.
After completing the regression task, the problem was reformulated as a binary classification task.
The continuous target Smoking_Prevalence was converted into two classes:
This produced a balanced dataset appropriate for training classification models.
The distribution of the new target variable was examined:
Since the classes were balanced, accuracy remained a meaningful metric.
However, precision, recall, and F1-score were also evaluated to gain deeper insight.
Three classification models were trained using the engineered features:
All models were trained using the same train/test split to ensure fair comparison.
| Model | Accuracy | Precision | Recall | F1 Score |
|---|---|---|---|---|
| Logistic Regression | ~0.50–0.60 | Moderate | Lower | Moderate |
| Random Forest Classifier | Highest | Highest | Highest | Best Overall |
| Gradient Boosting Class. | Close to RF | Slightly lower | Slightly lower | Strong |
The confusion matrices clearly show that Random Forest makes the fewest classification errors.
The Random Forest Classifier was selected as the winning model because it:
Both winning models were uploaded to the Hugging Face repository:
winning_model.pkl (Gradient Boosting)winning_classifier.pkl (Random Forest)