Downloads ยท 30 days
0
KalsusEvening/Melbourne_Housing_Regression
Melbourne_Housing_Regression is a machine learning model from KalsusEvening. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for sklearn. The card lists the license as mit.
Downloads ยท 30 days
0
Access
Public
Updated Dec 9, 2025
Repo size
2.2 MB
Likes
0
Public
Click a slice to open those files.
.ipynb6.5 MB ยท 59%
From the Hugging Face model README
(https://youtu.be/N3SE29PIr7g)
This project builds a complete machine learning pipeline to predict Melbourne housing prices using both regression (exact price) and classification (price category) models.
| Dataset | Melbourne Housing Snapshot (Kaggle) |
| Original Size | 13,580 properties, 21 features |
| Final Size | 11,139 properties (82% retained) |
| Target | Price |
| Step | Action | Impact |
|---|---|---|
| Missing Values | Dropped BuildingArea (47% missing), YearBuilt (40% missing) | - |
| Imputation | Car (median), CouncilArea (mode) | ~1,200 rows |
| Outliers | Removed using IQR method | 2,441 rows |
| Final | 11,139 rows retained | 18% removed |

Statistics: Mean $1.12M | Median $976K | Range $131K - $3.42M

| Type | Count | Mean Price |
|---|---|---|
| House | 9,055 (81%) | $1,203,259 |
| Townhouse | 952 (9%) | $936,054 |
| Unit | 1,132 (10%) | $640,529 |
Finding: Houses cost $560K more than units on average.

| Distance | Avg Price |
|---|---|
| 0-5 km | $1,361K |
| 10-15 km | $1,046K |
| 30+ km | $597K |
Finding: Every 5km from CBD reduces price by ~$100-150K. Correlation: -0.31

Finding: Southern Metropolitan commands 3.5ร premium over Western Victoria.

Top Correlations with Price:
| Metric | Value |
|---|---|
| Algorithm | Linear Regression |
| Features | 7 numeric |
| Rยฒ Score | 0.4048 |
| MAE | $323,527 |
| RMSE | $425,453 |
Interpretation: Model explains only 40% of price variance. Significant room for improvement through feature engineering.
| Category | Features | Count |
|---|---|---|
| Original Numeric | Rooms, Distance, Bathroom, etc. | 7 |
| One-Hot Encoded | Type, Method, Regionname | 16 |
| Derived Features | Ratios, indicators, bins | 7 |
| Cluster Features | Labels + distances to centroids | 8 |
| Total | 43 |
| Feature | Purpose |
|---|---|
| Rooms_per_Bathroom | Property efficiency |
| Total_Spaces | Overall size indicator |
| Land_per_Room | Land generosity |
| Is_Inner_City | Location premium flag |
| Luxury_Score | Amenity indicator |

We used the Elbow Method and Silhouette Score to determine k=4 clusters.

| Cluster | Profile | Avg Price | Avg Distance | Avg Rooms |
|---|---|---|---|---|
| 0 | Compact Inner Units | $835K | 8.0 km | 2.0 |
| 1 | Premium Family Estates | $1.48M | 12.9 km | 4.2 |
| 2 | Outer Suburban Affordable | $998K | 13.7 km | 3.0 |
| 3 | Inner City Houses | $1.18M | 7.9 km | 3.1 |
Key Insight: Two distinct pricing drivers discovered:

| Model | Rยฒ Score | MAE | Improvement |
|---|---|---|---|
| Baseline Linear Reg | 0.4048 | $323,527 | - |
| Improved Linear Reg | 0.6302 | $244,654 | +55.7% |
| Random Forest | 0.7752 | $178,455 | +91.5% |
| Gradient Boosting | 0.7900 | $172,891 | +95.1% |

Top 5 Most Important Features:
| Rank | Feature | Importance |
|---|---|---|
| 1 | Regionname_Southern Metropolitan | 0.242 |
| 2 | Distance | 0.172 |
| 3 | Type_h (House) | 0.137 |
| 4 | Dist_to_Cluster_0 | 0.099 |
| 5 | Landsize | 0.062 |
Key Insights:
| Metric | Value |
|---|---|
| Rยฒ Score | 0.7900 |
| MAE | $172,891 |
| RMSE | $252,728 |
| Improvement over Baseline | +95.1% |
Why Gradient Boosting Won:
Saved as: regression_model_gradient_boosting.pkl
We converted continuous Price into 3 balanced categories using quantile binning:

| Class | Price Range | Count | Percentage |
|---|---|---|---|
| Low | < $800,000 | 3,593 | 32.3% |
| Medium | $800K - $1.24M | 3,759 | 33.7% |
| High | > $1.24M | 3,787 | 34.0% |
Balance: Imbalance ratio of 1.05 - classes are well balanced.
For housing price prediction, Precision is more important:
| Error Type | Meaning | Consequence |
|---|---|---|
| False Positive | Predict High, actually Low | Buyer overpays significantly |
| False Negative | Predict Low, actually High | Seller underprices |
Conclusion: False Positives are worse for buyers - prioritize Precision.

| Model | Accuracy |
|---|---|
| Logistic Regression | 71.1% |
| Random Forest | 77.1% |
| Gradient Boosting | 78.9% |
| Class | Precision | Recall | F1-Score |
|---|---|---|---|
| Low | 0.85 | 0.86 | 0.85 |
| Medium | 0.69 | 0.70 | 0.70 |
| High | 0.83 | 0.81 | 0.82 |
Observations:
Saved as: classification_model_gradient_boosting.pkl
| File | Description | Size |
|---|---|---|
regression_model_gradient_boosting.pkl | Regression model (Rยฒ=0.79) | 419 KB |
classification_model_gradient_boosting.pkl | Classification model (78.9%) | 1.15 MB |
scaler.pkl | Regression StandardScaler | 2.32 KB |
classification_scaler.pkl | Classification StandardScaler | 2.32 KB |
feature_names.pkl | 43 feature names | 762 B |
Assignment_2_....ipynb | Complete Jupyter notebook | 6.48 MB |
| Task | Baseline | Final Model | Improvement |
|---|---|---|---|
| Regression Rยฒ | 0.4048 | 0.7900 | +95.1% |
| Regression MAE | $323,527 | $172,891 | -46.6% |
| Classification Accuracy | - | 78.9% | - |
| Features Used | 7 | 43 | +36 |
This project was completed with the assistance of Claude (Anthropic) as a coding and learning partner.
Why I used Claude:
What I learned through this process:
All code was executed, tested, and validated by me in Google Colab. The final analysis, interpretations, and conclusions are my own understanding of the results.
David Wilfand
Assignment #2: Classification, Regression, Clustering, Evaluation