Downloads · 30 days
13
4% of all-time downloads
Dunateo/roberta-cwe-classifier-kelemia
roberta-cwe-classifier-kelemia is a text classification model from Dunateo. Use it when you need a label for a piece of text. The card lists the license as mit.
This model is a fine-tuned version of RoBERTa for classifying Common Weakness Enumeration (CWE) vulnerabilities.
Downloads · 30 days
13
4% of all-time downloads
All-time downloads
311
Public
Parameters
125M
499 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors499 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of RoBERTa for classifying Common Weakness Enumeration (CWE) vulnerabilities.
Try now the v0.2 : Dunateo/roberta-cwe-classifier-kelemia-v0.2
This model is intended for classifying software vulnerabilities according to the CWE standard. It should be used as part of a broader security analysis process and not as a standalone solution for identifying vulnerabilities.
Here's an example of how to use this model for inference:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
# Load model and tokenizer
model_name = "Dunateo/roberta-cwe-classifier-kelemia"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
model.eval()
# Prepare input text
text = "The application stores sensitive user data in plaintext."
# Tokenize and prepare input
inputs = tokenizer(text, return_tensors="pt", truncation=True, padding=True, max_length=512)
# Perform inference
with torch.no_grad():
outputs = model(**inputs)
# Get prediction
probabilities = torch.nn.functional.softmax(outputs.logits, dim=-1)
predicted_class = torch.argmax(probabilities, dim=-1).item()
print(f"Predicted CWE class: {predicted_class}")
print(f"Confidence: {probabilities[predicted_class].item():.4f}")
This model uses the following mapping for CWE classes:
{
"0": "CWE-79",
"1": "CWE-89",
...
}
import json
from huggingface_hub import hf_hub_download
label_dict_file = hf_hub_download(repo_id="Dunateo/roberta-cwe-classifier-kelemia", filename="label_dict.json")
with open(label_dict_file, 'r') as f:
label_dict = json.load(f)
id2label = {v: k for k, v in label_dict.items()}
print(f"Label : {id2label[predicted_class]}")
| Epoch | Training Loss | Validation Loss |
|---|---|---|
| 1.0 | 4.822 | 4.639444828 |
| 2.0 | 3.6549 | 3.355055332 |
| 3.0 | 3.0617 | 2.821094036 |
The model shows consistent improvement over the training period:
This model should be used responsibly as part of a comprehensive security strategy. It should not be relied upon as the sole method for identifying or classifying vulnerabilities. False positives and negatives are possible, and results should be verified by security professionals.
For more details on the CWE standard, please visit Common Weakness Enumeration.
My report on this : Fine-tuning blogpost.