Downloads · 30 days
12
18% of all-time downloads
Aliph0th/logtheus-ml-large
logtheus-ml-large is a token classification model from Aliph0th. Use it when you need labels on individual words, such as names. It is set up for transformers.
A fine-tuned BERT model for extracting canonical attributes from log lines using token classification (NER-style task). Created for my course work
Downloads · 30 days
12
18% of all-time downloads
All-time downloads
66
Public
Parameters
334M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
A fine-tuned BERT model for extracting canonical attributes from log lines using token classification (NER-style task). Created for my course work
This model is based on bert-large-uncased and trained to perform token classification on log messages. It extracts structured attributes (service, level, event, error_code, user_id, ip, etc.) from unstructured log text using a BIO tagging scheme.
Use case: Convert raw log lines into canonical, structured key-value pairs for downstream analysis, alerting, or aggregation.
bert-large-uncasedThe model extracts attributes from these canonical fields:
| Field | Description |
|---|---|
| service | Application or service name (e.g., "auth", "api") |
| level | Log level (e.g., "info", "error", "warn") |
| timestamp | Timestamp or date reference |
| environment | Deployment environment (e.g., "prod", "staging") |
| event | Event type or action (e.g., "login", "request") |
| error_message | Human-readable error message |
| status_code | HTTP or service status code |
| duration | Duration |
| ip | IP address (client or server) |
| method | HTTP method (GET, POST, etc.) |
| path | URL path or resource path |
| useragent | User-Agent header |
| hostname | Server hostname |
pip install transformers torch
from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch
model_name = "Aliph0th/logtheus-ml"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(model_name)
text = "[auth] failed login for user 123 from 10.1.2.3 code=E401"
# Tokenize and forward pass
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=256)
outputs = model(**inputs)
logits = outputs.logits
# Get predicted label IDs
predicted_ids = torch.argmax(logits, dim=-1)
# Map back to label names
id2label = model.config.id2label
predictions = [[id2label[int(p)] for p in pred] for pred in predicted_ids]
print(predictions)
This returns a structured JSON object with:
attributes: High-confidence extractions (dict of canonical_field → value)low_confidence_attributes: Below-threshold extractionsattribute_confidence: Per-field confidence scoresmessage: Original log textconfidence: Overall prediction confidence (0-1)model_version: Model version stringTraining data in JSONL format with character-offset annotations:
{"id":"1","text":"[auth] failed login for user 123 from 10.1.2.3","entities":[{"start":1,"end":5,"label":"service"},{"start":28,"end":32,"label":"user_id"},{"start":38,"end":46,"label":"ip"}]}
Used dataset -Aliph0th/logtheus-ml-ds
Fields:
text: Raw log line (string)entities: List of entity annotations
start, end: Character-level offsets in text (0-indexed)label: Canonical field name# 1. Prepare raw log files (deduplicate, split train/val)
python scripts/process_data.py data/annotated/ --p 0.8
# 2. Train model
python training/train_token_classifier.py \
--train-file data/train.jsonl \
--val-file data/val.jsonl \
--output-dir artifacts/model_v1 \
--base-model bert-base-uncased \
--epochs 5 \
--batch-size 16
Hyperparameters:
For issues, questions, or contributions, please visit: