Downloads · 30 days
0
alperozyyurt/phishurl-detection
phishurl-detection is a tabular classification model from alperozyyurt. Use it for the tabular classification task on the model card, and read the license before you ship it in a product. It is set up for scikit-learn. The card lists the license as mit.
PhishURL Detection classifies URLs as legitimate or phishing/malicious using handcrafted URL features and multiple machine learning and deep learning models.
Downloads · 30 days
0
Access
Public
Updated Jun 9, 2026
Repo size
15.5 MB
Likes
0
Public
Click a slice to open those files.
.h510.1 MB · 63%
From the Hugging Face model README
PhishURL Detection classifies URLs as legitimate or phishing/malicious using handcrafted URL features and multiple machine learning and deep learning models.
Label convention:
0: legitimate / safe1: phishing / maliciousThe project includes classical ML and neural models trained on URL-derived features:
| Model | Test Accuracy | F1 Score | AUC |
|---|---|---|---|
| Random Forest | 0.9640 | 0.9640 | 0.9931 |
| XGBoost | 0.9587 | 0.9587 | 0.9935 |
| CNN | 0.9587 | 0.9587 | 0.9935 |
| Decision Tree | 0.9560 | 0.9560 | 0.9857 |
| ANN | 0.9547 | 0.9546 | 0.9920 |
| LightGBM | 0.9541 | 0.9541 | 0.9921 |
| DNN | 0.9219 | 0.9215 | 0.9175 |
The best reported model is Random Forest by test accuracy. XGBoost and CNN have the highest reported AUC among the included models.
The feature extractor creates URL-based signals including:
This model is intended for:
It should not be used as the only production security control. Real systems should combine model predictions with browser reputation, DNS intelligence, domain age, certificate metadata, sandboxing, and human review.
If this project helps your work, cite the repository:
@software{ozyurt_phishurl_detection_2026,
author = {Ozyurt, Alper},
title = {PhishURL Detection},
year = {2026},
url = {https://github.com/alperozyyurt4/phishurl}
}
Code and model packaging are released under the MIT License. Dataset redistribution rights must be verified separately before publishing the full dataset.