Downloads · 30 days
0
odedf2001/movies_metadata.csv
movies_metadata.csv is a tabular classification model from odedf2001. Use it for the tabular classification task on the model card, and read the license before you ship it in a product. It is set up for sklearn.
This project builds a complete machine learning workflow using real movie metadata. It includes data cleaning, exploratory data analysis (EDA), feature engineering, clustering, visualization, regression models, classi…
Downloads · 30 days
0
Access
Public
Updated Dec 8, 2025
Repo size
34.6 MB
Likes
0
Public
Click a slice to open those files.
.csv34.4 MB · 94%
From the Hugging Face model README
This project builds a complete machine learning workflow using real movie metadata.
It includes data cleaning, exploratory data analysis (EDA), feature engineering, clustering, visualization, regression models, classification models — and full performance evaluation.
Before any modeling, I asked a few basic questions about the dataset:
1️⃣ What is the relationship between budget and revenue?
2️⃣ Is there a strong relationship between runtime and revenue?
3️⃣ What are the most common original languages in the dataset?
These EDA steps helped build intuition before moving into modeling.
We test multiple regression models (Linear, Random Forest, Gradient Boosting) and evaluate how well different features explain revenue.
We explore the importance of:
We convert revenue into a balanced binary target and apply classification models.
We use K-Means + PCA to explore hidden groups, outliers, and natural segmentation of movies.
movies_metadata.csv (from Kaggle)revenue (continuous)budget, revenue, runtime, popularity to numeric.release_date as a datetime.budget == 0revenue == 0runtime == 0This produced a smaller but more reliable dataset.
Key insights:
Budget vs Revenue

Runtime vs Revenue

Original Language Distribution

These findings motivated the next steps: building a simple baseline model and then adding smarter features.
Build a simple baseline model that predicts movie revenue using only a few basic features:
budgetruntimevote_averagevote_countUsing only the basic features:
📌 Interpretation:
This baseline serves as a reference point before introducing engineered features.
To improve model performance, several new features were engineered:
profit = revenue - budgetprofit_ratio = profit / budgetoverview_length = length of the movie overview textrelease_year = year extracted from release_datedecade = grouped release year by decade (e.g., 1980, 1990, 2000)adult converted from "True"/"False" to 1/0.original_language and status encoded using One-Hot Encoding (with drop_first=True to avoid dummy variable trap).Used StandardScaler to standardize numeric columns:
budget, runtime, vote_average, vote_count,popularity, profit, profit_ratio, overview_lengthEach feature was transformed to have:
budget, runtime, vote_average, vote_count, popularity, profitn_clusters=4.cluster_group — each movie assigned to one of 4 clusters.Rough interpretation of clusters:
cluster_features to reduce dimensionality.pca1 and pca2 for each movie.cluster_group.This allowed visual inspection of:

Computed:
distance_to_centroid for each movie = Euclidean distance between the movie and its cluster center.Interpretation:
This feature was later used as an additional signal for modeling.

Use the engineered features + clustering-based features to improve regression performance.
Included:
budget, runtime, vote_average, vote_count, popularityprofit, profit_ratio, overview_length, release_year, decadecluster_group, distance_to_centroidoriginal_language_... and status_...| Model | MAE | RMSE | R² |
|---|---|---|---|
| Linear Regression | ~0 (leakage) | ~0 | 1.00 |
| Random Forest | 1,964,109 | 7,414,303 | 0.9975 |
| Gradient Boosting | 2,255,268 | 5,199,504 | 0.9988 |
📌 Note:
profit are directly derived from revenue).🔥 Gradient Boosting Regressor
Instead of predicting the exact revenue, we converted the problem to a binary classification task:
Class 1 (high revenue): 2687
Class 0 (low revenue): 2682
### 📊 Classification Results
#### Logistic Regression
- Accuracy: **0.977**
- Precision: **0.984**
- Recall: **0.968**
- F1: **0.976**
#### Random Forest
- Accuracy: **0.986**
- Precision: **0.988**
- Recall: **0.982**
- F1: **0.985**
#### Gradient Boosting Classifier
- Accuracy: **0.990**
- Precision: **0.990**
- Recall: **0.990**
- F1: **0.990**
---
## 🏆 Classification Winner
🔥 **Gradient Boosting Classifier**
- Highest accuracy
- Balanced precision & recall
- Best overall performance
---
## 📌 Tools Used
- Python
- pandas / numpy
- scikit-learn
- seaborn / matplotlib
- Google Colab
---
## 🎯 Final Summary
This project demonstrates a complete machine learning workflow:
- Data preprocessing
- Feature engineering
- K-Means clustering
- PCA visualization
- Regression models
- Classification models
- Full evaluation and comparison
The strongest model in both regression and classification tasks was **Gradient Boosting**, delivering state-of-the-art performance.
---
🎥 Watch the full project here: