Downloads · 30 days
262
63% of all-time downloads
rehan-ml/scamshield-scam-detector
scamshield-scam-detector is a text classification model from rehan-ml. Use it when you need a label for a piece of text. The card lists the license as mit.
A fine-tuned DistilBERT model that classifies text as scam or safe — built to run fully on-device (browser/edge), as the detection engine behind the ScamShield Chrome extension.
Downloads · 30 days
262
63% of all-time downloads
All-time downloads
416
Public
Parameters
67M
268 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors268 MB · 100%
From the Hugging Face model README
A fine-tuned DistilBERT model that classifies text as scam or safe — built to run fully on-device (browser/edge), as the detection engine behind the ScamShield Chrome extension.
This model detects scam, phishing, and fraudulent job/internship messages from plain text — SMS messages, emails, job offer letters, and similar written content. It was fine-tuned from distilbert-base-uncased for binary sequence classification (0 = safe, 1 = scam).
It's designed to be small and fast enough to run entirely client-side (in a browser via ONNX + Transformers.js, or on edge devices), so no user text needs to be sent to a server for scam detection.
distilbert-base-uncasedNot intended for: legal/compliance decisions, moderating content at scale without human review, or as a sole determinant of fraud — see Limitations below.
Combined from three sources:
| Metric | Score |
|---|---|
| Accuracy | 98.5% |
| Precision | 0.944 |
| Recall | 0.833 |
| F1 | 0.885 |
This model went through 4 iterations, which is worth documenting honestly since it shapes how the model should be used:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("rehan-ml/scamshield-scam-detector")
model = AutoModelForSequenceClassification.from_pretrained("rehan-ml/scamshield-scam-detector")
text = "Congratulations! You've been selected. Pay a refundable registration fee of $50 to confirm your position."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
outputs = model(**inputs)
scam_probability = torch.softmax(outputs.logits, dim=1)[0][1].item()
print(f"Scam probability: {scam_probability:.4f}")
Built by Rehan Raza for OSDHack 2026.