Downloads · 30 days
8
22% of all-time downloads
Rajith014/bert-phishing-detector
bert-phishing-detector is a text classification model from Rajith014. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
Fine-tuned bert-base-uncased for binary classification of email/message text as phishing or legitimate.
Downloads · 30 days
8
22% of all-time downloads
All-time downloads
37
Public
Parameters
109M
438 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
Fine-tuned bert-base-uncased for
binary classification of email/message text as phishing or legitimate.
| id | label |
|---|---|
| 0 | legitimate |
| 1 | phishing |
Not intended as a standalone production email security gateway. Always keep a human in the loop for high-stakes decisions, and combine with URL reputation, sender authentication (SPF/DKIM/DMARC) and other signals.
from transformers import pipeline
clf = pipeline("text-classification", model="<your-username>/bert-phishing-detector")
clf("Verify your account now at http://secure-login.example.ru")
# [{'label': 'phishing', 'score': 0.99}]
bert-base-uncasedTraining code: src/train.py in the project repository.
Fine-tuned for 3 epochs on ~65.7k emails and evaluated on an 8,207-email held-out test split:
| Metric | Score |
|---|---|
| Accuracy | 0.9948 |
| Precision | 0.9956 |
| Recall | 0.9944 |
| F1 | 0.9950 |
These scores are very high partly because the training data (the Kaggle phishing-email collection) is cleanly separable. Expect lower performance on noisier, real-world email — validate on your own data before relying on it.
data/sample_emails.csv is a tiny illustrative sample. Train on a
large, representative corpus before relying on any metric.Bring your own labeled corpus with a text column and an integer label
column (0 = legitimate, 1 = phishing). Commonly used public sources include the
Nazario phishing corpus, the Enron email dataset (legitimate mail), SpamAssassin
public corpus, and various Kaggle phishing-email datasets. Check each dataset's
license before redistribution.
This model classifies text and can make mistakes. Do not use it to automatically delete mail or take irreversible actions without human review.