Downloads · 30 days
0
asafmak/fifa_overall_predictor
fifa_overall_predictor is a machine learning model from asafmak. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<video controls width="100%" <source src="https://huggingface.co/asafmak/fifaoverallpredictor/resolve/main/Presentation-video-(1).mp4" type="video/mp4" </video
Downloads · 30 days
0
Access
Public
Updated Dec 8, 2025
Repo size
160 MB
Likes
0
Public
Click a slice to open those files.
.pkl97.2 MB · 82%
From the Hugging Face model README
This project analyzes FIFA player performance data and builds machine-learning models to predict and classify player ability.
The workflow includes:
Source: Kaggle – FIFA Players Dataset
Size: ~18,000 players
Features: Player attributes including age, height, weight, overall rating, potential, technical skills (passing, shooting, dribbling), physical abilities, positions, preferred foot, and market value/wage. Research Question
Which player attributes best predict a player’s overall rating, and how accurately can a model estimate this rating?
Main steps:



Purpose: Understand the data structure and detect patterns before modeling.

Insight: There are some differences in overall rating across positions, but the gaps are relatively small, and most positions fall within a similar overall range.

Insight: Younger players tend to have slightly higher potential, though the relationship is weak and there is wide variation across ages.

Insight: Defenders and goalkeepers tend to be taller, while attacking players and midfielders show more variance.

Insight: Defenders and goalkeepers tend to be taller, while attacking players and midfielders show more variance.

Insight: Right-footed and left-footed players have nearly identical distributions — preferred foot does not significantly impact player rating.
Goal: Predict a player’s overall rating using only the cleaned numerical features.

This plot shows how close model predictions are to the true ratings.

This shows how the prediction errors are distributed and whether the model is biased.

This highlights which numerical features most influence the model’s prediction.
The baseline linear regression model performs strongly, explaining over 91% of the variance in player ratings.
This provides a reliable starting point before adding engineered features or more advanced models.
These features capture broader skill profiles that individual columns cannot.
We compared multiple values of k using Silhouette Score and the Elbow Method.
Final choice: k = 4, as it provides the best balance between structure and interpretability.


Description: The PCA plot shows that players form four distinct clusters based on their abilities.
Each cluster represents a different skill profile, which is why Cluster ID is a meaningful new feature for the model.

I attempted to improve the Linear Regression model using the new engineered feature set, but its performance remained unchanged.
Linear Regression struggles to capture complex non-linear relationships in the data, so the additional features did not significantly improve its accuracy.
After this evaluation, I trained two additional models — Random Forest and Gradient Boosting — using the same improved features.
These models outperformed Linear Regression, with Random Forest achieving the best overall results.
| Model | MAE | RMSE | R² |
|---|---|---|---|
| Improved Linear Regression | 1.60 | 2.05 | 0.916 |
| Random Forest | 0.73 | 1.09 | 0.976 |
| Gradient Boosting | 0.86 | 1.19 | 0.972 |
Random Forest won because it captured non-linear patterns better and achieved the lowest errors with the highest R².
Converted the continuous target (overall rating) into three balanced classes using quantile binning:

To convert the regression task into classification, we trained three different classifiers (using the engineered features and the 3-class target):
I evaluated the performance of all three classification models using the same test set to ensure a fair and consistent comparison.
We assessed each model using Accuracy, Precision, Recall, and F1-Score, and also generated confusion matrices to better understand the types of errors each model makes.
| Model | Accuracy | Precision (macro) | Recall (macro) | F1 (macro) |
|---|---|---|---|---|
| Logistic Regression | 0.845 | 0.847 | 0.844 | 0.845 |
| Random Forest | 0.933 | 0.932 | 0.931 | 0.931 |
| Gradient Boosting | 0.944 | 0.943 | 0.942 | 0.943 |
Gradient Boosting achieved the best performance across accuracy, precision, recall, and F1-score.
This project demonstrates that combining clean data, feature engineering, and clustering significantly improves player rating prediction.
Tree-based models outperformed Linear Regression, with Gradient Boosting achieving the strongest classification performance.
Overall, the optimized feature set and model comparisons produced robust and reliable predictive results.