Downloads · 30 days
0
Leelu1002/spotify-popularity-predictor
spotify-popularity-predictor is a machine learning model from Leelu1002. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
For Video Presentation Please Click Here
Downloads · 30 days
0
Access
Public
Updated Dec 11, 2025
Repo size
35.6 MB
Likes
0
Public
Click a slice to open those files.
.ipynb6.3 MB · 100%
From the Hugging Face model README
For Video Presentation Please Click Here
This project predicts Spotify track popularity using audio features, metadata, feature engineering, clustering, regression models, and classification models.
It is based on the Spotify Tracks Dataset from Kaggle.
The goal of this project is to determine whether musical and audio characteristics can predict a track’s popularity (0–100). The task is first approached as a regression problem and later reframed as a binary classification problem.
Target variable: popularity
Type: continuous (0–100)
Source: Spotify Tracks Dataset, Kaggle
Size: ~114,000 rows, 20 features, 125+ genres
Includes:
Popularity is heavily right-skewed: most songs are minimally popular.
There are no strong linear correlations between audio features and popularity.
Several weak but consistent patterns appear:
There is a clear upward trend: songs with higher danceability tend to be more popular. While the relationship is not perfectly linear, the upper envelope rises consistently with danceability.2.Energy VS Popularity
Popular songs cluster around medium-to-high energy levels. Very low-energy tracks rarely achieve high popularity, showing a clear preference for energetic music.
3. Loudness VS Popularity
There is a visible positive trend: louder songs (closer to 0 dB) tend to achieve higher popularity. Quiet tracks rarely reach high popularity, reflecting modern production and streaming trends.



The data suggests that no single feature determines popularity.
Weak linear relationships indicate the need for non-linear models, feature engineering, and multi-feature interactions.
Genre features dominate model coefficients.
Audio trends include:

Among non-genre features:
StandardScaler applied to all numeric features.
PolynomialFeatures(degree=2) used to capture interactions and non-linear relationships.
Performed only for visualization (2 components).
No distinct clusters observed in PCA space.
K-Means (k=5) applied to scaled numeric features.
Added engineered features:
cluster_idcluster_distance
Clusters differ by energy, danceability, valence, acousticness, and avg popularity.
Three models were trained on the engineered dataset:
| Model | MAE | RMSE | R² |
|---|---|---|---|
| Enhanced Linear Regression | ~14.05 | ~19.03 | ~0.27 |
| Random Forest | ~15.97 | ~19.95 | ~0.20 |
| Gradient Boosting | ~17.02 | ~20.55 | ~0.15 |

Winner: Enhanced Linear Regression
Tree-based models underperformed due to:
Saved model: spotify_popularity_enhanced_linear_regression.pkl
A binary label was created using the training-set median popularity (35):
Classes were balanced (~50/50), so no resampling was needed.
Precision was prioritized over recall, since predicting a non-popular track as popular is more costly.
Three models were trained:
| Model | Accuracy |
|---|---|
| Logistic Regression | ~0.76 |
| Random Forest | ~0.75 |
| Gradient Boosting | ~0.72 |
Winner: Logistic Regression
It achieved the best precision-recall balance and the lowest misclassification bias.
Saved model: spotify_popularity_logistic_regression_classifier.pkl



Logistic Regression shows the best balance between precision and recall across both classes.
Install dependencies:
pip install -r requirements.txt
Run the notebook:
Spotify_Popularity_Classification,_Regression,_Clustering_Assignment_2.ipynb
The preprocessing pipeline includes:
These steps must be applied before loading any saved model.
project/
│── README.md
│── Leelu_Spotify_Popularity_Assignment_2.ipynb
│── spotify_popularity_enhanced_linear_regression.pkl
│── spotify_popularity_logistic_regression_classifier.pkl
This project builds a complete machine learning pipeline for predicting Spotify track popularity.
Through EDA, feature engineering, regression, and classification, the project demonstrates:
All final models require the full preprocessing pipeline to reproduce predictions.