Downloads · 30 days
4
15% of all-time downloads
muhaimin25/CORTEX-ATTACK-Classifier-v1
CORTEX-ATTACK-Classifier-v1 is a machine learning model from muhaimin25. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
4
15% of all-time downloads
All-time downloads
27
Public
Parameters
125M
499 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors499 MB · 99%
From the Hugging Face model README
tags:
text-classification
pytorch
transformers
cybersecurity
threat-intelligence
mitre-attack
bert
multi-label datasets:
tumeteor/Security-TTP-Mapping language:
en library_name: transformers pipeline_tag: text-classification license: apache-2.0 base_model: nanda-rani/TTPXHunter widget:
text: "The attacker used PowerShell to execute base64 encoded commands to download the payload." example_title: "PowerShell Execution"
text: "Lazarus Group malware has deleted files including suicide scripts to delete malware binaries." example_title: "Indicator Removal"
🧠 CORTEX-ATTACK-Classifier-v1
CORTEX-ATTACK-Classifier-v1 is a specialized BERT-based language model fine-tuned to automatically map unstructured cybersecurity threat reports (CTI) to specific MITRE ATT&CK® Enterprise Techniques.
It serves as the AI engine for the CORTEX Platform, designed to assist SOC analysts in rapidly triaging threat reports and generating D3FEND countermeasures.
🚀 Model Capabilities
Input: Unstructured text (Threat reports, logs, incident tickets, blog posts).
Output: A list of mapped MITRE ATT&CK Technique IDs (e.g., T1059.001) and their confidence scores.
Context Aware: Unlike keyword matching, this model uses semantic understanding to differentiate between mentioning a tool and using a tool for an attack.
💻 How to Use (Python)
This model comes with a custom metadata file (cortex_metadata.pkl) containing the exact label mappings and English technique names.
pip install transformers torch huggingface_hub
import torch import pickle from transformers import AutoTokenizer, AutoModelForSequenceClassification from huggingface_hub import hf_hub_download
MODEL_ID = "muhaimin25/CORTEX-ATTACK-Classifier-v1"
metadata_path = hf_hub_download(repo_id=MODEL_ID, filename="cortex_metadata.pkl") with open(metadata_path, 'rb') as f: metadata = pickle.load(f)
id2label = metadata['id2label'] ttp_names = metadata['ttp_names'] # English names (e.g., "Command and Scripting Interpreter")
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID) model = AutoModelForSequenceClassification.from_pretrained(MODEL_ID) model.eval()
def analyze_threat(text): inputs = tokenizer(text, return_tensors="pt", truncation=True, padding=True) with torch.no_grad(): outputs = model(**inputs)
# Apply Softmax for confidence scores
probs = torch.nn.Softmax(dim=0)(outputs.logits.squeeze())
# Get Top 3 Predictions
top_probs, top_indices = torch.topk(probs, 3)
results = []
for score, idx in zip(top_probs, top_indices):
if score < 0.10: continue # Filter low confidence
t_code = id2label.get(idx.item(), "Unknown")
name = ttp_names.get(t_code, "")
results.append(f"{t_code}: {name} ({score:.1%} Confidence)")
return results
text = "The attacker used PowerShell to execute base64 encoded commands." print(analyze_threat(text))
📊 Training Details
This model is a fine-tuned version of nanda-rani/TTPXHunter (based on RoBERTa/SecureBERT), specialized for the CORTEX environment.
Training Data
Dataset: tumeteor/Security-TTP-Mapping
Size: ~20,000 labeled sentences from real-world CTI reports.
Task: Multi-Label Text Classification (193+ Classes).
Training Procedure
Infrastructure: Fine-tuned on NVIDIA T4 GPUs via Google Colab.
Framework: PyTorch + Hugging Face Transformers.
Strategy: Strict Metadata Enforcement (Model configuration locked to a master dictionary of MITRE Techniques to prevent label misalignment).
Hyperparameters
Learning Rate: 2e-5
Batch Size: 16
Epochs: 2 (Early Stopping applied)
Weight Decay: 0.01
Optimizer: AdamW
📈 Evaluation & Metrics
The base model (TTPXHunter) and subsequent fine-tuning have demonstrated high performance in extracting TTPs from unstructured text.
Metric
Score (Approx.)
Description
F1-Score
~92-97%
Harmonic mean of precision and recall (based on TTPXHunter benchmarks).
Precision
High
Low false positive rate; reliable for automated triage.
Recall
High
Successfully identifies techniques even in verbose reports.
🛡️ Intended Use Case
This model is intended for:
Security Operations Centers (SOC): Automating the initial tagging of incident tickets.
Threat Intelligence: Extracting structured TTPs from blog posts and PDF reports.
Purple Teaming: Quickly mapping offensive actions to known techniques to validate defenses.
Limitations:
The model outputs probabilities based on text similarity. It should be verified by a human analyst.
It performs best on English technical text.
📚 Credits & Citations
Developed by: Mohammed Muhaimin
Base Model: TTPXHunter (Nanda Rani et al.)
Dataset: Security-TTP-Mapping (Tumeteor)
Framework: MITRE ATT&CK® (The MITRE Corporation)
Citation (BibTeX)
@misc{cortex-attack-classifier, author = {Muhaimin, Mohammed}, title = {CORTEX-ATTACK-Classifier-v1}, year = {2024}, publisher = {Hugging Face}, journal = {Hugging Face Model Hub}, howpublished = {\url{https://huggingface.co/muhaimin25/CORTEX-ATTACK-Classifier-v1}} }