Downloads ยท 30 days
0
mohitr7378/HousingPredictionApp
HousingPredictionApp is a machine learning model from mohitr7378. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This is a machine learning project where I built a model to predict the median house value of a district using the California Housing dataset.
Downloads ยท 30 days
0
Access
Public
Updated Aug 11, 2026
Repo size
145 MB
Likes
0
Public
Click a slice to open those files.
.pkl145 MB ยท 99%
From the Hugging Face model README
This is a machine learning project where I built a model to predict the median house value of a district using the California Housing dataset.
I wanted to understand the complete workflow of a regression problem instead of just training a model, so the project includes data splitting, preprocessing, model comparison, cross-validation, and finally using the trained model to make predictions.
The dataset contains information about different housing districts in California, including:
The target variable is median_house_value.
I used StratifiedShuffleSplit to divide the dataset into training and testing data.
I created an income_cat column based on median_income so that the income distribution remains similar in both sets.
For numerical features, I used:
For the categorical ocean_proximity feature, I used one-hot encoding.
I combined these using Scikit-learn's Pipeline and ColumnTransformer.
I tested three different regression models:
I used 10-fold cross-validation and RMSE to compare their performance.
| Model | Mean RMSE |
|---|---|
| Linear Regression | $69,070.52 |
| Decision Tree | $69,015.14 |
| Random Forest | $49,199.35 |
Lower RMSE means better performance, so Random Forest performed the best.
It reduced the RMSE by around 28.8% compared with Linear Regression, which made it the model I chose for the final prediction system.
After choosing Random Forest, I saved both the trained model and preprocessing pipeline using Joblib:
joblib.dump(model, "model.pkl")
joblib.dump(pipeline, "pipeline.pkl")
For new data, the saved pipeline preprocesses the input and the saved model generates the predicted house values.
The predictions are saved in output.csv.
California-Housing-Price-Prediction/
โ
โโโ housing.csv
โโโ input.csv
โโโ output.csv
โโโ ModelTest.py
โโโ main.py
โโโ model.pkl
โโโ pipeline.pkl
โโโ README.md
Install the required libraries:
pip install pandas numpy scikit-learn joblib
Run the project:
python main.py
The model will be trained if model.pkl doesn't already exist. Otherwise, the saved model will be used to make predictions from input.csv.
Through this project, I got hands-on experience with:
Some things I would like to add later:
Made by Mohit Raj AI & Data Science Student