Downloads · 30 days
1.3K
11% of all-time downloads
cisco-ai/SecureBERT2.0-code-vuln-detection
SecureBERT2.0-code-vuln-detection is a text classification model from cisco-ai. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
The ModernBERT Code Vulnerability Detection Model is a fine-tuned variant of SecureBERT 2.0, designed to detect potential vulnerabilities in source code. It leverages cybersecurity-aware representations learned by Sec…
Downloads · 30 days
1.3K
11% of all-time downloads
All-time downloads
12.2K
Public
Parameters
150M
598 MB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors598 MB · 99%
From the Hugging Face model README
The ModernBERT Code Vulnerability Detection Model is a fine-tuned variant of SecureBERT 2.0, designed to detect potential vulnerabilities in source code.
It leverages cybersecurity-aware representations learned by SecureBERT 2.0 and applies supervised fine-tuning for binary classification (vulnerable vs. non-vulnerable).
This model classifies source code snippets as either vulnerable or non-vulnerable using the ModernBERT architecture.
It is fine-tuned for code-level security analysis, extending the capabilities of SecureBERT 2.0.
ModernBertForSequenceClassificationCan be integrated into:
Users should use this model as an assistive tool, not as a replacement for expert manual code review.
Cross-validation with multiple tools is recommended before security-critical decisions.
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
# Path to the model
model_dir = "cisco-ai/SecureBERT2.0-code-vuln-detection"
# Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_dir)
model = AutoModelForSequenceClassification.from_pretrained(model_dir)
# Example input code snippet
example_code = """
static void FUNC_0(WmallDecodeCtx *VAR_0, int VAR_1, int VAR_2, int16_t VAR_3, int16_t VAR_4)
{
int16_t icoef;
int VAR_5 = VAR_0->cdlms[VAR_1][VAR_2].VAR_5;
int16_t range = 1 << (VAR_0->bits_per_sample - 1);
int VAR_6 = VAR_0->bits_per_sample > 16 ? 4 : 2;
if (VAR_3 > VAR_4) {
for (icoef = 0; icoef < VAR_0->cdlms[VAR_1][VAR_2].order; icoef++)
VAR_0->cdlms[VAR_1][VAR_2].coefs[icoef] +=
VAR_0->cdlms[VAR_1][VAR_2].lms_updates[icoef + VAR_5];
} else {
for (icoef = 0; icoef < VAR_0->cdlms[VAR_1][VAR_2].order; icoef++)
VAR_0->cdlms[VAR_1][VAR_2].coefs[icoef] -=
VAR_0->cdlms[VAR_1][VAR_2].lms_updates[icoef];
}
VAR_0->cdlms[VAR_1][VAR_2].VAR_5--;
}
"""
# Tokenize and run model
inputs = tokenizer(example_code, return_tensors="pt", truncation=True, padding=True)
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits
predicted_class = torch.argmax(logits, dim=-1).item()
print(f"Predicted class ID: {predicted_class}")
Internal validation split from annotated open-source vulnerability datasets.
Evaluated across:
| Model | Accuracy | F1 | Recall | Precision |
|---|---|---|---|---|
| CodeBERT | 0.627 | 0.372 | 0.241 | 0.821 |
| CyBERT | 0.459 | 0.630 | 1.000 | 0.459 |
| SecureBERT 2.0 | 0.655 | 0.616 | 0.602 | 0.630 |
SecureBERT 2.0 demonstrates the best overall balance of accuracy, F1, and precision among the compared models.
While CyBERT achieves the highest recall (detecting all vulnerabilities), it suffers from low precision, indicating many false positives.
Conversely, CodeBERT exhibits strong precision but poor recall, missing a large portion of true vulnerabilities.
SecureBERT 2.0 achieves more consistent and stable performance across all metrics, reflecting its stronger domain adaptation from cybersecurity-focused pretraining.
Carbon footprint can be estimated using the Machine Learning Impact Calculator.
BibTeX:
@article{aghaei2025securebert,
title={SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence},
author={Aghaei, Ehsan and Jain, Sarthak and Arun, Prashanth and Sambamoorthy, Arjun},
journal={arXiv preprint arXiv:2510.00240},
year={2025}
}
Cisco AI
For inquiries, please contact [email protected]