Downloads · 30 days
0
MingtongDu/ModernLogBERT
ModernLogBERT is a machine learning model from MingtongDu. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
ModernLogBERT is a log anomaly detection model based on the ModernBERT architecture, designed to process HDFS log data for identifying anomalies. This repository contains scripts for data preprocessing and model train…
Downloads · 30 days
0
Access
Public
Updated Jun 6, 2025
Repo size
3.1 GB
Likes
0
Public
Click a slice to open those files.
.log1.6 GB · 50%
From the Hugging Face model README
ModernLogBERT is a log anomaly detection model based on the ModernBERT architecture, designed to process HDFS log data for identifying anomalies. This repository contains scripts for data preprocessing and model training, leveraging a custom tokenizer and a pretrained ModernBERT model for efficient log sequence analysis.
The project includes 4 models (ModernLogbert_BC, ModernLogbert_BC_Logtokenizer, ModernLogbert_MLWP, ModernLogbert_MLWP_Logtokenizer), each model including two main scripts:
dataprocess.py: Preprocesses HDFS log files, extracts log messages, groups them by block ID, and tokenizes sequences for model input.train.py: Trains the LogBERT model using Masked Log Word Prediction (MLWP) and KL divergence tasks for anomaly detection.The model is built on the answerdotai/ModernBERT-base model, enhanced with modern architectural improvements like Rotary Positional Embeddings (RoPE) and Flash Attention for long-context processing.
huggingface)Install dependencies using:
pip install torch transformers pandas numpy scikit-learn
Clone the Repository:
git clone https://huggingface.co/MingtongDu/ModernLogBERT
cd ModernLogBERT
Prepare Data:
HDFS.log) and anomaly label file (e.g., anomaly_label.csv) in an accessible directory.dataprocess.py and train.py to match your environment.Tokenizer:
train.py.Pretrained Model:
ep165.pt). Update the CHECKPOINT_PATH in train.py to point to your checkpoint file.Run dataprocess.py to parse, tokenize, and save the processed datasets:
python dataprocess.py
This script:
HDFS.log to extract log messages and block IDs.anomaly_label.csv.train.pt, val.pt, and test.pt in the specified output directory.Run train.py to train the LogBERT model:
python train.py
This script:
EP3_froze21.pt.Key hyperparameters are defined at the top of each script:
dataprocess.py: MAX_LENGTH, TRAIN_RATIO, VAL_RATIO, TEST_RATIO, PRINT_SAMPLES, RANDOM_SEED.train.py: BATCH_SIZE, NUM_EPOCHS, LEARNING_RATE, NUM_LAYERS_TO_FREEZE, DROPOUT_RATE, MASK_PROB, ALPHA, G, R.Modify these in the scripts to adjust preprocessing or training behavior.
dataprocess.py: Script for log parsing, label loading, grouping, tokenization, and data saving.train.py: Script for loading data, defining the LogBERT model, and training with anomaly detection metrics.output/: Directory containing processed datasets (train.pt, val.pt, test.pt) and the saved tokenizer.saved_model/: Directory for saving the trained model checkpoint.If you use ModernLogBERT, please cite the ModernBERT model:
ModernLogBERT: Mingtong Du and Haoran Li. https://huggingface.co/MingtongDu/ModernLogBERT
This project is licensed under the MIT License. See the LICENSE file for details.