Downloads · 30 days
0
Yuvalos/Yuval-Malka-GTD-Regression-Model
Yuval-Malka-GTD-Regression-Model is a machine learning model from Yuvalos. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Assignment 2 — Yuval Malka Classification, Regression, Clustering & Evaluation Global Terrorism Database (GTD)
Downloads · 30 days
0
Access
Public
Updated Dec 11, 2025
Repo size
1.3 MB
Likes
0
Public
Click a slice to open those files.
.ipynb2.7 MB · 67%
From the Hugging Face model README
Assignment 2 — Yuval Malka Classification, Regression, Clustering & Evaluation Global Terrorism Database (GTD)
Video Presentation: https://youtu.be/6yeC_aRLvTU Part 1: Dataset Overview
For this assignment, I selected the Global Terrorism Database (GTD) from Kaggle. The dataset contains approximately 181,000 terrorism incidents recorded worldwide from 1970 to 2017. It includes 135 features, combining both numeric variables (casualty counts, coordinates, temporal attributes) and categorical descriptors (country, region, attack type, weapon type, target type, etc.).
The main regression question explored in this project was: Can we predict the number of wounded (nwound) in a terrorism event based on event characteristics such as location, weapon type, and attack type?
sample data preview

Part 2: Exploratory Data Analysis (EDA)
The EDA process focused on understanding the distribution of casualties, identifying missing values, detecting outliers, and exploring differences between attack types, regions, and weapon categories.
Key steps included:
Handling missing values: casualty-related NaNs were replaced with zeros; categorical missing values were set to “Unknown”.
Outlier detection: extreme casualty values were retained because they represent real, important events.
Converting latitude and longitude to numeric and removing invalid geographic entries.
Producing descriptive statistics for all numerical fields.
Examining correlations between important variables.
Visualizing severity across regions, weapon types, attack types, and time.
boxplot of nwound






Summary of EDA: The dataset is highly skewed, with most events causing zero or few casualties. Severity varies significantly across regions and attack types, indicating meaningful patterns for modeling. Geographic clustering shows concentrated hotspots of high-severity events. These observations guided feature engineering decisions in later stages.
Additional Research Questions
Several analytical questions were explored using visualizations:
How did the number of terrorist attacks change over time?

Which attack types are most common?

Which regions suffer the most average casualties?

Which weapon types tend to cause the highest severity?

Do suicide attacks cause more casualties compared to non-suicide attacks?

Are high-severity events geographically clustered?

These visual insights further demonstrated distinct severity patterns across event categories.
Part 3: Baseline Regression Model (Explanation)
The baseline Linear Regression model required for Part 3 is fully implemented later in Part 5, after feature engineering and PCA. This ordering ensures that the baseline is trained on a clean and meaningful feature set, rather than on raw, unprocessed data.
To keep the assignment structure complete, the Part 3 baseline code is shown below but not executed, since running it before preprocessing would produce inconsistent results. The evaluated baseline model — along with metrics (MAE, MSE, RMSE, R²) and comparisons — appears in Part 5 and satisfies all requirements of this section.
Part 4: Feature Engineering
This stage introduced several new features and transformations aimed at improving model performance.
Engineered numeric features included:
total_casualties (nkill + nwound)
severity_index (log-transformed total casualties)
attack_age (years since attack)
is_successful (binary outcome)
suicide_weapon_interaction (interaction term)
Categorical features were one-hot encoded (attack type, target type, weapon type, country, region), producing over 255 dummy variables.
Numeric features were standardized using StandardScaler.
Dimensionality reduction was then applied using PCA to condense the dataset into 50 principal components, reducing noise and improving model efficiency.
Unsupervised Learning: KMeans clustering (k=5) was performed on numeric features, and the resulting cluster_label was added as an engineered feature. PCA visualization of clusters revealed meaningful separation, suggesting that the clusters captured underlying event behavior patterns.

Summary: The combination of engineered features, scaling, encoding, PCA, and clustering produced a rich and informative feature set to support both regression and classification tasks.
Part 5: Regression Models — Training and Evaluation
Three regression models were trained using the engineered dataset:
Linear Regression (SGDRegressor)
Random Forest Regressor
Gradient Boosting Regressor
Before training, PCA-reduced features (50 components) were sampled to 20,000 rows for efficiency. An 80/20 train-test split was used.
Regression Results:
model comparison table

Summary of Results:
The Linear Regression baseline performed poorly because it could not capture nonlinear event patterns.
Random Forest significantly improved error metrics.
Gradient Boosting achieved the overall best performance with the lowest RMSE and highest R² score.
Winning regression model: Gradient Boosting Regressor This model was exported and uploaded as winning_regression_model.pkl.

Part 6: Exporting the Winning Regression Model
The Gradient Boosting model was saved using joblib and uploaded to the HuggingFace repository. File: winning_regression_model.pkl
Part 7: Regression-to-Classification Conversion
To convert the regression target (nwound) into categories, quantile-based thresholds were applied:
Class 0: Low severity (bottom 33%)
Class 1: Medium severity (33–66%)
Class 2: High severity (top 33%)
Due to the heavy skew toward zero casualties, quantile boundaries were computed manually.
balance distribution:

The dataset is imbalanced, particularly for the medium-severity class. Therefore metrics like recall and macro-F1 are more meaningful than accuracy.
Part 8: Classification Models — Training and Evaluation
Using the PCA-reduced feature matrix, three different classification models were trained:
Logistic Regression
Random Forest Classifier
HistGradientBoostingClassifier
Each model was evaluated using classification reports and confusion matrices.
logistic regression confusion matrix random forest confusion matrix HistGradientBoosting confusion matrix



Summary of Findings:
Logistic Regression performed strongly on low- and high-severity classes but struggled with medium severity.
Random Forest produced more errors across classes and was less stable.
HistGradientBoostingClassifier performed the most consistently and accurately across all classes, especially reducing errors in the medium and high categories.
Winning classification model: HistGradientBoostingClassifier This model was exported as winning_classification_model.pkl.
Part 9: Presentation Video
The video presentation includes:
An introduction to the dataset and research question
Key EDA insights and visualizations
The full feature engineering process
Description of PCA and clustering
Regression model training and comparison
Classification model training and evaluation
Lessons learned and reflections
Part 10: Repository Contents
This repository includes:
README file
Jupyter Notebook (.ipynb)
winning_regression_model.pkl
winning_classification_model.pkl
Presentation video link