Downloads · 30 days
118
7% of all-time downloads
gneeraj/deeppass2-bert
deeppass2-bert is a token classification model from gneeraj. Use it when you need labels on individual words, such as names. The card lists the license as apache-2.0.
DeepPass2 is a fine-tuned version of xlm-roberta-base specifically designed for detecting passwords and secrets in documents through token classification. Unlike traditional regex-based approaches, this model understa…
Downloads · 30 days
118
7% of all-time downloads
All-time downloads
1.6K
Public
Parameters
277M
1.1 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
DeepPass2 is a fine-tuned version of xlm-roberta-base specifically designed for detecting passwords and secrets in documents through token classification. Unlike traditional regex-based approaches, this model understands context to identify both structured tokens (API keys, JWTs) and free-form passwords.
Developed by: Neeraj Gupta (SpecterOps)
Model type: Token Classification (Sequence Labeling)
Base model: xlm-roberta-base
Language(s): English
License: [Same as base model]
Fine-tuned with: LoRA (Low-Rank Adaptation) through Unsloth
Blog post: What's Your Secret?: Secret Scanning by DeepPass2
LoraConfig(
task_type=TaskType.TOKEN_CLS,
r=64, # Rank
lora_alpha=128, # Scaling parameter
lora_dropout=0.05, # Dropout probability
bias="none",
target_modules=["query", "key", "value", "dense"]
)
This model is the BERT based model used in the DeepPass2 blog.
0: Non-credential token1: Credential/password token"Your account has been created with username: {user} and password: {pass}"
# Preprocessing
- Tokenization with offset mapping
- Label generation based on credential spans
- Padding to max_length with truncation
# Fine-tuning
- LoRA adapters applied to attention layers
- Binary cross-entropy loss
- Token-level classification head
| Metric | Score |
|---|---|
| Strict Accuracy | 86.67% |
| Overlap Accuracy | 97.72% |
| Metric | Count/Rate |
|---|---|
| True Positives | 1,201 |
| True Negatives | 1,112 |
| False Positives | 49 (3.9%) |
| False Negatives | 138 |
| Overlap True Positives | 456 |
| Recall | 89.7% |
pip install transformers torch
from transformers import AutoModelForTokenClassification, AutoTokenizer
import torch
# Load model and tokenizer
model_name = "path/to/deeppass2-xlm-roberta"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForTokenClassification.from_pretrained(model_name)
# Classify tokens
def detect_passwords(text):
inputs = tokenizer(text, return_tensors="pt", truncation=True, padding=True)
with torch.no_grad():
outputs = model(**inputs)
predictions = torch.argmax(outputs.logits, dim=-1)
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
# Extract password tokens
password_tokens = [
token for token, label in zip(tokens, predictions[0])
if label == 1
]
return password_tokens
For production use, integrate with the full DeepPass2 pipeline:
See the DeepPass2 repository for complete implementation.
@software{gupta2025deeppass2,
author = {Gupta, Neeraj},
title = {DeepPass2: Fine-tuned XLM-RoBERTa for Secret Detection},
year = {2025},
organization = {SpecterOps},
url = {https://huggingface.co/deeppass2-bert},
note = {Blog: \url{https://specterops.io/blog/2025/07/31/whats-your-secret-secret-scanning-by-deeppass2/}}
}
For questions or issues, please open an issue on the GitHub repository