Downloads · 30 days
0
shiraBASH/flight-price-random-forest
flight-price-random-forest is a machine learning model from shiraBASH. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
0
Access
Public
Updated Dec 11, 2025
Repo size
2.5 GB
Likes
0
Public
Click a slice to open those files.
.pkl2.4 GB · 98%
From the Hugging Face model README
Author: Shira Bash
project walkthrough presentation:
Loom Video
This project aims to analyze and predict flight prices using multiple machine learning approaches.
The assignment includes:
The goal is to understand what drives flight prices, build effective predictive models, and communicate insights clearly.
The dataset contains detailed information about commercial flights, including airline, departure/arrival times, number of stops, duration, and price. After cleaning (removing duplicates, dropping an unnamed column, and removing outliers in two columns using IQR), the final dataset shape is:
airline, source_city, departure_time, stops, arrival_time
destination_city, class, duration, days_left
Target (Regression): price
Conclusion: Ticket prices are highly right-skewed: most tickets are relatively cheap, with a small number of very expensive flights forming a long right tail of the distribution.
Conclusion: Flights with one or more stops tend to be much more expensive than non-stop flights, and 1-stop itineraries are the priciest on average – showing that the number of stops is strongly related to ticket price.
Conclusion: Longer flights generally have higher ticket prices, but the scatter is large, so duration alone cannot explain price; other features such as class, airline and number of stops also play an important role.
1. How does the average ticket price vary by month?
Conclusion: The average ticket price drops noticeably from February to March. This suggests a seasonal pattern where flights in March tend to be cheaper, possibly due to lower demand or post-holiday travel trends.
2. How do ticket prices vary across travel class and number of stops?
Conclusion: Business-class tickets are consistently more expensive than Economy across all stop categories. Prices also increase sharply with additional stops, especially for Business class. This indicates a strong interaction between travel class and number of stops in determining ticket price.
3. What is the relationship between flight duration and ticket price for non-stop flights?
Conclusion: There is a positive trend—longer non-stop flights tend to have higher ticket prices. However, the large scatter shows that duration alone cannot fully explain price variation, meaning additional factors like class, route, and airline also significantly influence ticket cost.
4. How is the price distribution different between Economy and Business class?
Conclusion: Economy tickets are heavily concentrated in the low-price range, while Business-class tickets appear at much higher prices with a wider spread. This clear separation shows that travel class is one of the strongest predictors of ticket price.
These insights were later connected to model performance and feature importance.
Using cleaned + encoded baseline features (11 features).
Train set:
Test set:
The baseline model performs strongly, showing that the engineered features are meaningful.
Conclusion: The baseline linear regression captures the general upward trend between true and predicted prices, but many points deviate from the ideal diagonal line. This indicates that the model struggles with non-linear patterns and underfits in several price ranges, especially for very cheap and very expensive tickets.
To improve model performance, I added several engineered features on top of the baseline numeric features.
Using K-Means (k = 4), I created:
cluster_id – assigned cluster label
cluster_distance – distance to cluster centroid
PCA 2D visualization of clusters

These features help the models capture non-linear patterns and interactions in the data.
| Model | MAE | RMSE | R² |
|---|---|---|---|
| Random Forest | best | best | highest |
| Gradient Boosting | slightly worse | slightly worse | lower |
Random Forest was the winning regression model.

Conclusion: The Random Forest model identifies class as by far the most important feature for predicting ticket prices. Other features such as duration_in_min, route_index, and airline_index contribute much less but still show meaningful influence. This confirms that travel class is the strongest driver of ticket pricing in the dataset.
In addition to predicting the exact ticket price, I also reframed the problem as a 3-class classification task.
The goal here is to classify each flight into a price level (cheap / medium / expensive) based on the numeric target price.
price_class targetTo create the classification labels I transformed the continuous price variable into three classes using quantile binning on the training set only:
This strategy has two advantages:
The resulting discrete labels are stored in a new target column called price_class, which is used throughout the classification part.
After creating the labels, I checked the class distribution in both the train and test sets.
Each of the three classes (0 – cheap, 1 – medium, 2 – expensive) contains roughly one third of the samples in both splits (≈ 32–34% per class).
This means the classes are reasonably balanced and:

Using the engineered features from the regression part, I trained several classification models to predict the new target price_class (0 = cheap, 1 = medium, 2 = expensive).
From a user perspective, the most important group is the cheap tickets (class 0):
In terms of error types:
I trained three different classifiers on the same feature matrix (X_train_fe, X_test_fe) and the new target price_class:
All models were evaluated on the same train/test split using:
classification_report (precision, recall, F1, support) for train and test setsOn the test set, the models achieved the following results:
| Model | Accuracy (test) | Macro F1 (test) |
|---|---|---|
| Logistic Regression | 0.794 | 0.795 |
| Gradient Boosting | 0.869 | 0.869 |
| Random Forest | 0.953 | 0.952 |
Because the classes are balanced, these metrics can be compared fairly across models.
Logistic Regression

Gradient Boosting

Random Forest

The Random Forest Classifier is the winning model for the classification task because:
Loom Video - project walkthrough presentation
Freshly_cleaned.csv - dataset used for the analysis
Assignment#2-Classification,Regression,Clustering,Evaluation-Shira Bash.ipynb – full Python notebook
README.md – summary of results, questions, and insights
flight_price_random_forest_classifier.pkl - 1 pickle model for classification
flight_price_random_forest.pkl - 1 pickle model for regression
This project demonstrates the full ML lifecycle: