Downloads · 30 days
0
galsaar/chess-game-length-predictor
chess-game-length-predictor is a machine learning model from galsaar. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
▶️ Project Presentation Video: https://www.loom.com/share/7e7a5a47fec8469f9b457a910fc52c0d
Downloads · 30 days
0
Access
Public
Updated Dec 11, 2025
Repo size
287 MB
Likes
1
Public
Click a slice to open those files.
.pkl287 MB · 98%
From the Hugging Face model README
▶️ Project Presentation Video: https://www.loom.com/share/7e7a5a47fec8469f9b457a910fc52c0d
This repository contains the final trained model from an end-to-end machine learning project analyzing ~20K chess games sourced from Lichess. The goal was to predict the length of a chess game (in number of half-moves) by classifying each game into one of three categories:
The project includes exploratory data analysis, feature engineering, categorical encoding, model training, and exporting the final model as a pickle file.
Chess games vary greatly in structure and duration. Some games end quickly due to opening traps or mistakes, while others develop into long positional struggles.
This model predicts a game’s length category using information available before or shortly after the opening phase, including:
The final exported artifact is a trained Random Forest Classifier.
The dataset contains approximately 19,800 chess games with:
The target variable (turns_class) was created by binning number of turns into three quantile-based classes.
Key insights included:
Some openings consistently lead to longer or shorter games.
Higher-rated matchups tend to produce longer, higher-quality games.
Games ending by resignation or timeout skew shorter.
Hundreds of unique openings required careful preprocessing and encoding.
Engineered features include:
opening_moves was grouped into four categories: very_short, short, medium, long.
rating_diff = abs(white_rating - black_rating)rating_avg = (white_rating + black_rating) / 2Boolean feature marking whether a game ended in a draw.
KMeans clustering was used to group openings with similar behavior.
Categorical variables were transformed using one-hot encoding, with filtering to avoid high dimensionality.
Three models were evaluated:
Baseline model. Accuracy: ~46–47%.
Best-performing model overall. Accuracy: ~47–48%.
Slightly below Random Forest.
Random Forest achieved the strongest performance due to:
This repository hosts the exported model:
random_forest_model.pkl
| File | Description |
|---|---|
random_forest_model.pkl | Final trained model |
README.md | Project documentation |
.gitattributes | Managed by HuggingFace |
Copy_of_Assignment_2_Classification,_Regression,_Clustering,_Evaluation.ipynb | Google Colab |
Thanks to Lichess for providing open game data and to HuggingFace for model hosting.