Downloads · 30 days
187
84% of all-time downloads
Podric/prowl-secret-encoder
prowl-secret-encoder is a text classification model from Podric. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
<p align="center"<img src="https://huggingface.co/Podric/prowl-secret-encoder/resolve/main/logo.png" width="360" alt="Prowl"</p
Downloads · 30 days
187
84% of all-time downloads
All-time downloads
222
Public
Parameters
135M
541 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors541 MB · 99%
From the Hugging Face model README
A multilingual text classifier that scores whether a fragment of text (a line of code, a config value, a Jira comment, a log line, a chat message) contains a leaked credential. It is stage 3 of Prowl, a high-precision secret scanner, and handles the free-form and multilingual tail that regex and the linear model miss.
This model is not a standalone scanner. It is one stage of an ensemble; on its own it is a recall booster, not a precision oracle. To scan a repo, use Prowl.
softmax(...)[1] is P(text contains a secret).P ≥ 0.90, a threshold calibrated on a held-out validation split to
precision ≥ 0.95 (value-disjoint from the benchmark, no leakage).Prowl combines three stages by union: a Go regex/checksum/entropy cascade, a char+word TF-IDF logistic regression, and this encoder. Adding the encoder to the other two, measured on ProwlBench (3,843 leakage-safe cases):
| Configuration | Precision | Recall | F1 |
|---|---|---|---|
| cascade ∪ LR | 0.974 | 0.848 | 0.909 |
| + this encoder | 0.971 | 0.970 | 0.970 |
The encoder lifts recall by 0.12 at a 0.003 precision cost, and reaches recall 1.00 on German, French, Spanish, and Russian prose passwords - the free-form, multilingual tail that regex and the linear model miss.
Standalone recall at a fixed precision of 0.97, by channel:
| code | Jira | Confluence | log |
|---|---|---|---|
| 0.984 | 0.996 | 0.998 | 0.947 |
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tok = AutoTokenizer.from_pretrained("Podric/prowl-secret-encoder")
model = AutoModelForSequenceClassification.from_pretrained("Podric/prowl-secret-encoder").eval()
def secret_score(text: str) -> float:
enc = tok(text, return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
return torch.softmax(model(**enc).logits, -1)[0, 1].item()
secret_score('DB_PASSWORD = "tR4!nf0rce-2026-prod"') # high
secret_score("the deployment finished without errors") # low
# fire at >= 0.90
secret_score(text) on a few inputs (the model fires at P ≥ 0.90):
| P(secret) | fires | input |
|---|---|---|
| 1.000 | yes | DB_PASSWORD = "tR4!nf0rce-2026-prod" |
| 1.000 | yes | Ihr neues Passwort lautet F4QPE91sc6iN ... (German prose) |
| 1.000 | yes | Пароль от прод-сервера: Zx9!kLmN2k24qP ... (Russian prose) |
| 0.001 | no | token = os.environ["SERVICE_TOKEN"] (env reference) |
| 0.002 | no | Развёртывание завершилось без ошибок ... (Russian log) |
| 0.000 | no | def calculate_total(items): return sum(...) (benign code) |
| 0.000 | no | API_KEY=your_api_key_here (placeholder) |
It fires on real passwords, including free-form prose in German and Russian that has no fixed prefix for a regex to anchor, and stays near zero on benign code, logs, env-var references, and placeholders.
distilbert-base-multilingual-cased (104 languages, 6 layers, 134M params).@misc{prowl_secret_encoder,
title = {Prowl secret encoder},
author = {Prowl},
year = {2026},
howpublished = {Hugging Face},
url = {https://huggingface.co/Podric/prowl-secret-encoder}
}
Noncommercial use only (CC BY-NC 4.0). Not for use in commercial products. Fine-tuned from
distilbert-base-multilingual-cased (Apache-2.0). See the
Prowl repository and the dataset card for data provenance and
source licenses.