Downloads · 30 days
9
30% of all-time downloads
ctokx/cti-attack-mapper-modernbert
cti-attack-mapper-modernbert is a text classification model from ctokx. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
Maps sentences from cyber threat intelligence reports to MITRE ATT&CK technique IDs. 49 techniques, multi-label, sentence level.
Downloads · 30 days
9
30% of all-time downloads
All-time downloads
30
Public
Parameters
150M
599 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors599 MB · 99%
From the Hugging Face model README
Maps sentences from cyber threat intelligence reports to MITRE ATT&CK technique IDs. 49 techniques, multi-label, sentence level.
Read this before you trust the numbers. On the leak-free split, a fine-tuned ModernBERT scores 0.445 macro-F1. A plain TF-IDF baseline scores 0.454. The transformer does not beat it; they are tied inside the noise of a 151-report corpus. What does beat both is a blend of the two: 0.474 macro-F1 (5-fold mean, standard deviation 0.022), about 0.04 ahead of each base model at the same threshold setting, and ahead in all five folds. One caveat: against TF-IDF at its own best threshold setting the blend's lead shrinks to 0.02 and is no longer outside the noise. Every number here is written by a script in the repo. See Results.
Defensive use only. It labels adversary behavior that has already been written up in public threat reports. It produces no offensive capability.
There are already several ATT&CK classifiers on the Hub. This one exists for three reasons those usually leave out.
scripts/04_report.py from JSON that the training and eval scripts produce.
None of it is typed by hand.from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
model_id = "ctokx/cti-attack-mapper-modernbert"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id).eval()
text = "The implant establishes persistence by creating a scheduled task that runs at logon."
with torch.no_grad():
probs = torch.sigmoid(model(**tok(text, return_tensors="pt", truncation=True)).logits)[0]
for i, p in enumerate(probs):
if p > 0.5:
print(model.config.id2label[i], round(float(p), 3))
# T1053.005 0.94
Or use the repo's wrapper, which loads the tuned per-class thresholds and resolves technique names:
from cti_attack.predict import AttackMapper
mapper = AttackMapper("models/modernbert__document")
mapper.predict("The dropper base64-encodes its configuration before writing it to disk.")
# [Prediction(T1027, 'Obfuscated Files or Information', 0.912)]
Per-class thresholds tuned on dev. Macro-F1 is the headline metric. It weights all 49 techniques equally, so a handful of common ones cannot cover for a weak long tail.
| Model | doc macro-F1 | doc micro-F1 | random macro-F1 | random micro-F1 | inflation |
|---|---|---|---|---|---|
| Frequency prior | 0.000 | 0.000 | 0.000 | 0.000 | n/a |
| ATT&CK keyword match | 0.167 | 0.170 | 0.171 | 0.153 | +2.4% |
| TF-IDF + one-vs-rest LR | 0.454 | 0.500 | 0.506 | 0.545 | +11.5% |
| ModernBERT-base (fine-tuned) | 0.445 | 0.478 | 0.482 | 0.540 | +8.3% |
| SecureBERT (fine-tuned) | 0.369 | 0.252 | 0.415 | 0.373 | +12.5% |
| Model | doc macro-F1 | doc micro-F1 | random macro-F1 | random micro-F1 | inflation |
|---|---|---|---|---|---|
| Frequency prior | 0.000 | 0.000 | 0.000 | 0.000 | n/a |
| ATT&CK keyword match | 0.167 | 0.170 | 0.171 | 0.153 | +2.4% |
| TF-IDF + one-vs-rest LR | 0.449 | 0.472 | 0.515 | 0.552 | +14.7% |
| ModernBERT-base (fine-tuned) | 0.425 | 0.494 | 0.504 | 0.557 | +18.6% |
| SecureBERT (fine-tuned) | 0.352 | 0.425 | 0.388 | 0.470 | +10.2% |
One 70/15/15 split of a 151-document corpus is one draw, so the gap between models can be noise. To check it, here is 5-fold cross-validation grouped by document. Each report lands in exactly one test fold. The blend weight and thresholds are tuned on each fold's own dev set, never on its test set.
| Model | doc macro-F1 (5-fold CV, mean +/- std) |
|---|---|
| TF-IDF + LR (per-class) | 0.4326 +/- 0.0349 |
| TF-IDF + LR (global) | 0.4521 +/- 0.0420 |
| ModernBERT (per-class) | 0.4263 +/- 0.0256 |
| Ensemble TF-IDF+ModernBERT (per-class) | 0.4738 +/- 0.0223 |
Paired per-fold test, per-class thresholds on both sides:
On the original single 70/15/15 document split, the same ensemble (70% TF-IDF, 30% ModernBERT) scores 0.4965 macro-F1, against 0.4536 for TF-IDF and 0.4454 for ModernBERT.
How to read this: the ensemble has the highest mean and the lowest variance of everything tested. Compared against each base model at the same threshold setting, it wins in all five folds. The one comparison it does not clearly win is against TF-IDF at its own best threshold setting, where the lead is small enough to be noise over five folds. That row is in the table on purpose.
| Technique | Name | Support | P | R | F1 |
|---|---|---|---|---|---|
T1027 | Obfuscated Files or Information | 102 | 0.57 | 0.54 | 0.55 |
T1140 | Deobfuscate/Decode Files or Information | 68 | 0.73 | 0.84 | 0.78 |
T1105 | Ingress Tool Transfer | 57 | 0.51 | 0.44 | 0.47 |
T1059.003 | Command and Scripting Interpreter: Windows … | 51 | 0.39 | 0.59 | 0.47 |
T1055 | Process Injection | 40 | 0.68 | 0.53 | 0.59 |
T1106 | Native API | 34 | 0.64 | 0.47 | 0.54 |
T1047 | Windows Management Instrumentation | 28 | 0.69 | 0.64 | 0.67 |
T1053.005 | Scheduled Task/Job: Scheduled Task | 26 | 0.71 | 0.92 | 0.80 |
T1562.001 | Impair Defenses: Disable or Modify Tools ⚠️revoked | 26 | 0.92 | 0.46 | 0.62 |
T1574.002 | Hijack Execution Flow: DLL Side-Loading ⚠️revoked | 26 | 0.83 | 0.38 | 0.53 |
T1082 | System Information Discovery | 24 | 0.83 | 0.21 | 0.33 |
T1078 | Valid Accounts | 23 | 0.23 | 0.48 | 0.31 |
| Bucket | Techniques | Mean F1 |
|---|---|---|
| head (>=20 test examples) | 15 | 0.570 |
| tail (<20 test examples) | 34 | 0.391 |
random column is the number you would have reported by accident. It
runs about 12 percent higher for the same models on the same data. If you are
comparing against work that split at sentence level, that column is the
comparable one, and it is not the true one.In scope. A triage aid for threat-intelligence and detection-engineering work:
Out of scope.
T1562.001 and
T1574.002 were valid when the TRAM corpus was annotated and MITRE has since
revoked them. Map them forward before comparing output to current ATT&CK.T1027 has 678 training instances; the rarest retained
techniques have about 20.| Base model | answerdotai/ModernBERT-base (149M) |
| Objective | Multi-label BCE, per-class pos_weight, capped at 50 |
| Max sequence length | 256 tokens |
| Batch size | 16, gradient accumulation 2 (effective 32) |
| Learning rate | 3e-5, linear schedule, 10% warmup |
| Weight decay | 0.01 (excluding bias and norm parameters) |
| Epochs | 6, best checkpoint by dev macro-F1 (epoch 5) |
| Precision | bf16 autocast |
| Hardware | 1x RTX 4060 Laptop, 8 GB |
| Peak VRAM | ~5.5 GB |
| Wall clock | ~30 min per split scheme |
| Seed | 20260802 |
The pos_weight is doing real work. 78 percent of sentences carry no label, and
without it the model learns to predict nothing.
Training loss reached 0.010 while dev macro-F1 peaked at 0.414. The model memorised the training set instead of generalising from it. With 2,684 labelled training sentences across 49 techniques, that is what you would expect, and it is the likely reason a single transformer does not pull ahead of the linear baseline. A blend of the two does, which points at data, not architecture, as the limit.
See DATASET_CARD.md. In short:
random split exposes.T1557.001 dropped: it appears in exactly one document and cannot be split
leak-freepip install -r requirements.txt
python scripts/reproduce_all.py # ModernBERT only, ~40 min on an RTX 4060
python scripts/reproduce_all.py --all-models # adds DeBERTa-v3 and SecureBERT
Individual steps:
python scripts/01_build_dataset.py # build both splits
python scripts/02_run_baselines.py # frequency, keyword, TF-IDF
python scripts/03_train.py --model modernbert --scheme document # fine-tune
python scripts/05_ensemble.py --scheme document # blend TF-IDF + ModernBERT
python scripts/06_cv.py --folds 5 # 5-fold document CV + error bars
python scripts/04_report.py # regenerate the tables above
pytest tests/ -q # invariants + smoke test
The published weights are the ModernBERT model. The blend is those weights plus
a TF-IDF and logistic regression model, which scripts/05_ensemble.py refits
from the included dataset in a few seconds, so the best-scoring setup reproduces
without shipping a second binary. The blend weight and every threshold are tuned
on dev, never on test.
The test suite checks the claims this card makes: that no document spans two splits, that all 49 techniques reach every split, that no duplicate sentences survive, and that the model loads and returns well-formed output.
Apache-2.0. Derived from MITRE CTID TRAM (Apache-2.0). Technique names come from MITRE ATT&CK STIX data under the ATT&CK Terms of Use.
ATT&CK® is a registered trademark of The MITRE Corporation. This project is not affiliated with, endorsed by, or sponsored by The MITRE Corporation.
See NOTICE for full attribution.
@misc{cti_attack_mapper,
title = {cti-attack-mapper: sentence-level MITRE ATT&CK classification
with leak-free evaluation},
author = {Varol Cagdas Tok},
year = {2026},
url = {https://huggingface.co/ctokx/cti-attack-mapper-modernbert},
note = {Derived from MITRE CTID TRAM, Apache-2.0}
}