Downloads · 30 days
0
SivakumarP/PhishingURLDetection
PhishingURLDetection is a machine learning model from SivakumarP. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
The objective of the project is to classify if an URL is phishing or not. This model repo contains the required encoders (for url,dom and tld), scaler (for digitcnt and ishttps) and the trained model (RandomForest Cla…
Downloads · 30 days
0
Access
Public
Updated Jul 12, 2026
Repo size
31.2 MB
Likes
0
Public
Click a slice to open those files.
.pkl31.2 MB · 100%
From the Hugging Face model README
The objective of the project is to classify if an URL is phishing or not. This model repo contains the required encoders (for url,dom and tld), scaler (for digit_cnt and is_https) and the trained model (RandomForest Classifier).
This project uses the URL-Phish dataset. The dataset was obtained from Kaggle, where it is available as Phishing URL Detection (111K URLs, 22 Features).
The dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0), which permits sharing, redistribution, and adaptation with appropriate credit.
Dataset citation<br>
Dam Minh, Linh; Tran Cong, Hung (2025).<br> URL-Phish: A Feature-Engineered Dataset for Phishing Detection.<br> Mendeley Data, V1.<br> DOI: https://doi.org/10.17632/65z9twcx3r.1<br>
Original data sources referenced by the dataset authors<br>
PhishTank – Community-driven phishing URL repository<br> Research Organization Registry (ROR) dataset – Source of trusted benign domain URLs<br>
Paper citation<br>
Dam Minh Linh, Tran Cong Hung, <br> A feature-engineered dataset of benign and phishing URLs for machine learning and large language models evaluation,<br> Data in Brief,<br> Volume 63,<br> 2025,<br> 112162,<br> ISSN 2352-3409,<br> https://doi.org/10.1016/j.dib.2025.112162.
Modifications: <br> The following preprocessing was applied to the original dataset:
Feature usage: <br> The final model was trained using a selected subset of the features; the remaining features were excluded at training time via feature selection, not by removing them from the stored datasets.
LICENSE