Downloads · 30 days
0
sassoftware/glitext-pii-large
glitext-pii-large is a token classification model from sassoftware. Use it when you need labels on individual words, such as names. It is set up for glitext. The card lists the license as other.
GLiNER-PII is inspired by the Gretel GLiNER PII/PHI models. Built on the GLiNER large-v2.1 base, it detects and classifies a broad range of Personally Identifiable Information (PII) and Protected Health Information (P…
Downloads · 30 days
0
Access
Public
Updated Jun 19, 2026
Repo size
3.6 GB
Likes
0
Public
Click a slice to open those files.
.onnx1.8 GB · 50%
From the Hugging Face model README
GLiNER-PII is inspired by the Gretel GLiNER PII/PHI models. Built on the GLiNER large-v2.1 base, it detects and classifies a broad range of Personally Identifiable Information (PII) and Protected Health Information (PHI) in structured and unstructured text. It is non-generative and produces span-level entity annotations with confidence scores across 55+ categories. This model was developed by NVIDIA.
This model is ready for commercial/non-commercial use. <br>
Use of this model is governed by the NVIDIA Open Model License Agreement. <br>
Global <br>
GLiNER-PII supports detection and redaction of sensitive information across regulated and enterprise scenarios.
Note: performance varies by domain, format, and threshold, so validation and human review are recommended for high‑stakes deployments. <br>
Hugging Face 10/28/2025 via https://huggingface.co/nvidia/gliner-pii <br>
Architecture Type: Transformer <br>
Network Architecture: GLiNER <br>
This model was developed based on urchade/gliner_large-v2.1 <br> Number of model parameters: 5.7 × 10^8 <br>
Input Type(s): Text <br> Input Format: UTF-8 string(s) <br> Input Parameters: One-Dimensional (1D) <br> Other Properties Related to Input: supports structured and unstructured text <br>
Output Type(s): Text <br> Output Format: String <br> Output Parameters: One-Dimensional (1D) <br> Other Properties Related to Output: List of dictionaries with keys {text, label, start, end, score} <br>
Runtime Engine(s):
Supported Hardware Microarchitecture Compatibility: <br>
Preferred/Supported Operating System(s):
The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment. <br>
Link: nvidia/nemotron-pii <br> Data Modality: Text <br> Text Training Data Size: ~100k records (~10^5, <1B tokens) <br> Data Collection Method: Synthetic <br> Labeling Method: Synthetic <br>
Properties: Synthetic persona-grounded dataset generated with NVIDIA NeMo Data Designer, spanning 50+ industries and 55+ entity types (U.S. and international formats). Includes both structured and unstructured records. Labels automatically injected during generation.
Data Collection Method: Hybrid: Automated, Human <br> Labeling Method: Hybrid: Automated, Human <br>
Evaluation Results <br> From the combined evaluation across Argilla, AI4Privacy, and Gretel PII datasets:
| Benchmark | Strict F1 |
|---|---|
| Argilla PII | 0.70 |
| AI4Privacy | 0.64 |
| nvidia/Nemotron-PII | 0.87 |
We evaluated the model using threshold=0.3. <br>
Acceleration Engine: PyTorch (via Hugging Face Transformers) <br> Test Hardware: NVIDIA A100 (Ampere, PCIe/SXM) <br>
First, make sure you have the gliner library installed:
pip install gliner
Now, let's try to find an email, SSN, and phone number in a messy block of text.
from gliner import GLiNER
# 1. Define our new text
text = "Hi support, I can't log in! My account username is 'johndoe88'. Every time I try, it says 'invalid credentials'. Please reset my password. You can reach me at (555) 123-4567 or [email protected]"
# 2. Define the labels we're hunting for.
labels = ["email", "phone_number", "user_name"]
# 3. Load the PII model
model = GLiNER.from_pretrained("nvidia/gliner-pii")
# 4. Run the prediction at given threshold
entities = model.predict_entities(text, labels, threshold=0.5)
Sample output:
[
{
"start": 52,
"end": 61,
"text": "johndoe88",
"label": "user_name",
"score": 0.99
},
{
"start": 159,
"end": 173,
"text": "(555) 123-4567",
"label": "phone_number",
"score": 0.99
},
{
"start": 177,
"end": 194,
"text": "[email protected]",
"label": "email",
"score": 0.99
}
]
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. <br>
For more detailed information on ethical considerations for this model, please see the Bias, Explainability, Safety & Security, and Privacy Subcards. <br>
Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here. <br>
This model is derived from nvidia/gliner-PII. See the upstream repository for the original safetensors weights, training data, and the full upstream model card.
ONNX weights added by SAS — converted from the upstream safetensors checkpoint.
File in this repo: model.onnx.
This repo is consumed by the SAS GLiText product. To download it onto a SAS GLiText server:
POST /v1/models/download?name=pii-large
To download and load into memory in one step:
PUT /v1/models?name=pii-large
Scanned with modelaudit v0.2.40 on 2026-04-27. 29/29 checks passed. Full results.
| File | Size | SHA-256 |
|---|---|---|
model.onnx | 1784.6 MB | e6f5db22bf3a2c06… |