Downloads · 30 days
0
danadvash/startupfunding
startupfunding is a machine learning model from danadvash. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<video controls width="700" <source src="https://huggingface.co/danadvash/startupfunding/resolve/main/Assignment%202%20-%20Dana%20Dvash%20Intro%20to%20Data%20Science%20(1).mp4" type="video/mp4" / Your browser does not…
Downloads · 30 days
0
Access
Public
Updated Dec 9, 2025
Repo size
57.2 MB
Likes
0
Public
Click a slice to open those files.
.mp417.4 MB · 83%
From the Hugging Face model README
This project analyzes startup funding data with the goal of understanding the drivers of startup success and building predictive models that estimate funding outcomes.
The project combines exploratory data analysis, feature engineering, clustering, and both regression and classification models to extract actionable insights for investors and analysts.
The dataset contains information about startups, including:
Basic preprocessing and cleaning steps were applied before modeling.
Q1: What factors are most associated with higher startup funding?
Startups with more funding rounds, longer operating history, and later-stage investments tend to receive higher total funding.
Q2: Can startups be meaningfully grouped using unsupervised learning?
Yes. Clustering revealed clear groups representing early-stage, growth-stage, and highly funded startups.
Q3: Does feature engineering improve model performance?
Yes. Aggregated funding features and cluster-based features significantly improved both regression accuracy and classification F1-scores.
Q4: Is reframing the problem as classification useful?
Absolutely. Classifying startups into high-funded vs. low-funded provides a practical screening tool for investment decision-making.
EDA focused on understanding the distribution of funding amounts, identifying skewness, and examining relationships between key variables.

Key insights:
Key feature engineering steps included:
K-Means clustering was applied on scaled features to create a new categorical feature representing startup funding profiles.

Clustering helped distinguish between early-stage, growth-stage, and mature startups.
A regression model was trained to predict total funding amount.
The trained regression model was exported as a pickle file.
The continuous funding target was converted into a binary classification problem using a median split:
This strategy ensured well-balanced classes.
Three different classification models were trained and evaluated using the same engineered features:
Model performance was evaluated using:

✅ Gradient Boosting was selected as the final classification model.
Reasons:
The trained classification model was exported as a pickle file.
notebook.ipynb – Full analysis and modeling pipelineREADME.md – Project documentationregression_model.pkl – Trained regression modelclassification_model.pkl – Winning classification model