Downloads · 30 days
873
3% of all-time downloads
cisco-ai/SecureBERT2.0-NER
SecureBERT2.0-NER is a token classification model from cisco-ai. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
The Secure Modern BERT NER Model is a fine-tuned transformer based on SecureBERT 2.0, designed for Named Entity Recognition (NER) in cybersecurity text.
Downloads · 30 days
873
3% of all-time downloads
All-time downloads
33.3K
Public
Parameters
150M
598 MB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors598 MB · 99%
From the Hugging Face model README
The Secure Modern BERT NER Model is a fine-tuned transformer based on SecureBERT 2.0, designed for Named Entity Recognition (NER) in cybersecurity text.
It extracts domain-specific entities such as Indicators, Malware, Organizations, Systems, and Vulnerabilities from unstructured data sources like threat reports, incident analyses, advisories, and blogs.
NER in cybersecurity enables:
| Entity | Description |
|---|---|
B-Indicator, I-Indicator | Indicators of Compromise (e.g., IPs, domains, hashes) |
B-Malware, I-Malware | Malware or exploit names |
B-Organization, I-Organization | Companies or groups mentioned |
B-System, I-System | Affected software or platforms |
B-Vulnerability, I-Vulnerability | Specific CVEs or flaw descriptions |
O | Outside token |
| Parameter | Value |
|---|---|
| Hidden size | 768 |
| Intermediate size | 1152 |
| Hidden layers | 22 |
| Attention heads | 12 |
| Max sequence length | 8192 |
| Vocabulary size | 50368 |
| Activation | GELU |
| Dropout | 0.0 (embedding, attention, MLP, classifier) |
This model can be integrated into:
| Aspect | Description |
|---|---|
| Purpose | Benchmark dataset for extracting cybersecurity entities from unstructured reports |
| Data Source | Curated threat intelligence documents emphasizing malware and system analysis |
| Annotation Methodology | Fully hand-labeled by domain experts |
| Entity Types | Malware, Indicator, System, Organization, Vulnerability |
| Size | 3.4k training samples + 717 test samples |
from transformers import AutoTokenizer, TFAutoModelForTokenClassification, pipeline
model_name = "cisco-ai/SecureBERT2.0-NER"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = TFAutoModelForTokenClassification.from_pretrained(model_name)
ner_pipeline = pipeline("ner", model=model, tokenizer=tokenizer)
text = "Stealc malware targets browser cookies and passwords."
entities = ner_pipeline(text)
print(entities)
The SecureBERT2.0-NER was fine-tuned for token-level classification on cybersecurity text using Cross Entropy Loss.
Training focused on accurately classifying entity boundaries and types across five cybersecurity-specific categories: Malware, Indicator, System, Organization, and Vulnerability.
The AdamW optimizer was used with a linear learning rate scheduler, and gradient clipping ensured stability during fine-tuning.
| Setting | Value |
|---|---|
| Objective | Token-wise Cross Entropy |
| Optimizer | AdamW |
| Learning Rate | 1e-5 |
| Weight Decay | 0.001 |
| Batch Size per GPU | 8 |
| Epochs | 20 |
| Max Sequence Length | 1024 |
| Gradient Clipping Norm | 1.0 |
| Scheduler | Linear |
| Mixed Precision | fp16 |
| Framework | TensorFlow / Transformers |
The model was fine-tuned on a cybersecurity-specific NER corpus, containing annotated threat intelligence reports, advisories, and technical documentation.
| Property | Description |
|---|---|
| Dataset Type | Manually annotated corpus |
| Language | English |
| Entity Types | Malware, Indicator, System, Organization, Vulnerability |
| Train Size | 3,400 samples |
| Test Size | 717 samples |
| Annotation Method | Expert hand-labeling for accuracy and consistency |
PreTrainedTokenizerFast tokenizer from SecureBERT 2.0.| Component | Description |
|---|---|
| GPUs Used | 8× NVIDIA A100 |
| Precision | Mixed precision (fp16) |
| Batch Size | 8 per GPU |
| Framework | Transformers (TensorFlow backend) |
The model converged after approximately 20 epochs, with loss stabilizing at a low level.
Validation metrics (F1, precision, recall) showed steady improvement from epoch 3 onward, confirming effective domain-specific adaptation.
Evaluation was conducted on a cybersecurity-specific NER benchmark corpus containing annotated threat reports, advisories, and incident analysis texts.
This benchmark includes five key entity types: Malware, Indicator, System, Organization, and Vulnerability.
The following metrics were used to assess model performance:
| Model | F1 | Recall | Precision |
|---|---|---|---|
| CyBERT | 0.351 | 0.281 | 0.467 |
| SecureBERT | 0.734 | 0.759 | 0.717 |
| SecureBERT 2.0 (Ours) | 0.945 | 0.965 | 0.927 |
The SecureBERT 2.0 NER model significantly outperforms both CyBERT and the original SecureBERT across all metrics.
This demonstrates that domain-adaptive pretraining and fine-tuning on cybersecurity corpora dramatically improves NER performance compared to general or earlier models.
@article{aghaei2025securebert,
title={SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence},
author={Aghaei, Ehsan and Jain, Sarthak and Arun, Prashanth and Sambamoorthy, Arjun},
journal={arXiv preprint arXiv:2510.00240},
year={2025}
}
Cisco AI
For inquiries, please contact [email protected]