Downloads · 30 days
2
9% of all-time downloads
daslabhq/pii-proxy
pii-proxy is a token classification model from daslabhq. Use it when you need labels on individual words, such as names. It is set up for gliner. The card lists the license as apache-2.0.
A GLiNER model fine-tuned for PII detection. GLiNER does zero-shot NER over arbitrary labels — you pass the entity types you care about at inference time, so there is no fixed label schema to work around.
Downloads · 30 days
2
9% of all-time downloads
All-time downloads
23
Public
Repo size
611 MB
Likes
0
Public
Click a slice to open those files.
.bin611 MB · 99%
From the Hugging Face model README
A GLiNER model fine-tuned for PII detection. GLiNER does zero-shot NER over arbitrary labels — you pass the entity types you care about at inference time, so there is no fixed label schema to work around.
Fine-tuned from urchade/gliner_small-v2.1 on the full Nemotron-PII dataset (~100k).
Held-out test set (100 examples), 24 fine-grained PII labels:
| Metric | F1 |
|---|---|
| Fine-grained | 92.8% |
| Coarse-grained | 94.4% |
NVIDIA's Nemotron-PII reference: fine 96.2%, coarse 96.7%.
from gliner import GLiNER
model = GLiNER.from_pretrained("daslabhq/pii-proxy")
labels = ["first_name", "last_name", "email", "phone_number", "ssn", "street_address"]
entities = model.predict_entities("Patient Marcus Weber, [email protected]", labels)
for e in entities:
print(e["text"], "->", e["label"])
urchade/gliner_small-v2.1