Downloads · 30 days
866
0% of all-time downloads
AungMoonLord/bert-log-anomaly-detection
bert-log-anomaly-detection is a text classification model from AungMoonLord. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
866
0% of all-time downloads
All-time downloads
670K
Public
Parameters
109M
438 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
Model Summary
bert-log-anomaly-detection is a BERT-based NLP model fine-tuned for single SQL transaction log anomaly detection.
The model classifies each database transaction log as either Normal or Anomaly, with the goal of supporting AI-powered fraud detection and cybersecurity monitoring systems.
This model was developed as part of the Samsung × KBTG Digital Fraud Cybersecurity Hackathon (Thailand) under the AI-Powered Fraud Detection & Prevention track.
This model analyzes individual SQL database transaction logs and detects abnormal patterns that may indicate fraudulent, malicious, or suspicious behavior.
Demo: Hackathon prototype
import torch
from transformers import BertForSequenceClassification, BertTokenizer
MODEL_PATH = "AungMoonLord/bert-log-anomaly-detection"
model = BertForSequenceClassification.from_pretrained(MODEL_PATH)
tokenizer = BertTokenizer.from_pretrained(MODEL_PATH)
model.eval()
# Perfom log preprocessing
def add_prefix_token(text): # log data must pass this code before training/inferencing
# clean log
text = text.replace("\t", " ")
text = text.strip()
# add token
if text[0].isalpha() or text[3].isalpha():
return "[SQL]\n" + text
else:
return "[LOG]\n" + text
def predict_log(log_text):
log_text = add_prefix_token(log_text)
inputs = tokenizer(
log_text,
return_tensors="pt",
truncation=True,
padding=True, # for cases when the inference contains more than 1 log, i.e., batch size > 1
max_length=128
)
with torch.no_grad():
logits = model(**inputs).logits
pred = torch.argmax(logits, dim=1).item()
prob = torch.softmax(logits, dim=-1).tolist()[0]
return "Normal" if pred == 1 else "Anomaly", prob
# Example 1
text1 = "SELECT * FROM users WHERE id = 1 OR 1=1"
print(predict_log(text1))
# Example 2
text2 = "2025-01-06 14:23:45 | User: anonymous | IP: 203.154.89.102 | Duration: 0.05s SELECT * FROM users WHERE username = 'admin' OR '1'='1' -- ' AND password = 'x'"
print(predict_log(text2))
# Example 3
text3 = "3051-06-22T07:20:02.296945Z 3 Query select e3mJKDCCY from 7Q8SpG8LLEWhrfpe4s5 where ph4d = 'a1S9hQa92uC1EAyJf2Y';"
print(predict_log(text3))
Multi-log sequence anomaly detection
Non-textual anomaly detection
SQL database transaction logs (1,611 samples) synthetically generated by ChatGPT, Qwen, DeepSeek, Grok, Gemini, and Claude
Each log labeled as either Normal or Anomaly
Data prepared for single-log classification
| Metric | Value |
|---|---|
| Accuracy | 0.8950 |
| Precision | 0.8580 |
| Recall | 0.9026 |
| F1-score | 0.8797 |
| Validation Loss | 0.3279 |
| Metric | Value |
|---|---|
| Accuracy | 0.6950 |
| Precision | 0.6639 |
| Recall | 0.7900 |
| F1-score | 0.7215 |
| Validation Loss | 0.6251 |
| Metric | Value |
|---|---|
| Accuracy | 0.7000 |
| Precision | 0.6613 |
| Recall | 0.8200 |
| F1-score | 0.7321 |
| Validation Loss | 0.6344 |
The model demonstrates strong anomaly detection capability with high recall, making it suitable for fraud detection and cybersecurity use cases.