Downloads · 30 days
77
10% of all-time downloads
Gaykar/PhishingDistilBERT
PhishingDistilBERT is a machine learning model from Gaykar. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc.
PhishingDistilBERT is a DistilBERT-based NLP model fine-tuned specifically for email understanding tasks, particularly phishing and suspicious email detection. The model introduces custom special tokens to explicitly…
Downloads · 30 days
77
10% of all-time downloads
All-time downloads
784
Public
Parameters
67M
268 MB on disk
Likes
6
Public
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
PhishingDistilBERT is a DistilBERT-based NLP model fine-tuned specifically for email understanding tasks, particularly phishing and suspicious email detection.
The model introduces custom special tokens to explicitly encode email structure such as subject, body, links, and phone numbers, making it more robust for email-based security applications.
It can be used both as:
This model is fine-tuned from distilbert-base-uncased on curated email datasets. During preprocessing, email-specific entities such as URLs and phone numbers are replaced with dedicated tokens, and the subject and body are explicitly separated using structural markers.
Special Tokens Used
[SSUB], [ESUB] – Start/End of Subject[SBODY], [EBODY] – Start/End of Body[LINK] – URLs[PHONE] – Phone numbersThese design choices help the model better learn semantic and structural patterns commonly found in phishing emails.
Users should carefully evaluate the model in their target environment before deployment.
from transformers import DistilBertTokenizerFast, DistilBertForSequenceClassification
import torch
import numpy as np
bert_path = "Gaykar/PhishingDistilBERT"
tokenizer = DistilBertTokenizerFast.from_pretrained(bert_path)
model = DistilBertForSequenceClassification.from_pretrained(bert_path)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)
model.eval()
def get_cls_embedding(text, model, tokenizer, device):
with torch.no_grad():
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
padding=True,
max_length=256
)
inputs = {k: v.to(device) for k, v in inputs.items()}
outputs = model.distilbert(**inputs)
cls_embedding = outputs.last_hidden_state[:, 0, :].squeeze().cpu().numpy()
return cls_embedding
text = "[SSUB] Urgent Account Alert [ESUB] [SBODY] Click [LINK] to verify your account. [EBODY]"
embedding = get_cls_embedding(text, model, tokenizer, device)
print("Embedding shape:", embedding.shape)
print("First 10 dimensions:", embedding[:10])
The model was trained using well-known phishing and email security datasets, including CEAS, combined with additional curated CSV sources.
Cleaned and merged multiple CSV datasets
Replaced:
[LINK][PHONE]Combined subject and body using structural tokens:
[SSUB], [ESUB], [SBODY], [EBODY]training_args = TrainingArguments(
output_dir="./distilbert_safe_suspicious",
eval_strategy="steps",
eval_steps=50,
save_strategy="steps",
save_steps=50,
save_total_limit=3,
load_best_model_at_end=True,
metric_for_best_model="eval_loss",
greater_is_better=False,
learning_rate=4e-5,
per_device_train_batch_size=16,
per_device_eval_batch_size=8,
num_train_epochs=4,
weight_decay=0.01,
logging_strategy="steps",
logging_steps=50,
seed=42,
)

Carbon emissions were not explicitly measured. Users may estimate emissions using the Machine Learning Impact Calculator if needed.
For questions, feedback, or research collaboration, please reach out via the Hugging Face model repository.