Downloads · 30 days
21
36% of all-time downloads
ManiKumarAdapala/distilbert-pii-ner
distilbert-pii-ner is a token classification model from ManiKumarAdapala. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as afl-3.0.
ADAPALA MANI KUMAR BHAT MITALI MAHENDRA CHELLAPPAN C ELLURU SAI GAGAN MD FAREED FAROOQUI
Downloads · 30 days
21
36% of all-time downloads
All-time downloads
59
Public
Parameters
65.3M
261 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors261 MB · 100%
From the Hugging Face model README
Finetuned Model & Project URL : https://drive.google.com/drive/folders/19ZP1RG_9Ms_kzsiLYBMrtrUbETej5dLq?usp=sharing
This project implements a PII (Personally Identifiable Information) detection and masking system using DistilBERT fine-tuned on the ai4privacy/pii-masking-200k dataset. The system exposes a Flask API for uploading text files and returning masked outputs.
.
.
├── app.py
├── design_document.docx
├── distilbert-ner
│ └── checkpoint-10880
│ ├── config.json
│ ├── model.safetensors
│ ├── tokenizer_config.json
│ ├── tokenizer.json
│ ├── trainer_state.json
│ └── training_args.bin
├── NER_Masking.pdf
├── NER_Masking.ipynb
├── readme.md
├── sample.txt
└── templates
└── index.html
https://drive.google.com/drive/folders/19ZP1RG_9Ms_kzsiLYBMrtrUbETej5dLq?usp=sharing
conda create -n pii python=3.10
conda activate pii
pip install gdown torch matplotlib transformers datasets seqeval scikit-learn seaborn numpy ipywidgets
To Train the Model run every cell in the file NER_Masking.ipynb
To run the Flask API:
python app.py
The server will start locally (default: http://127.0.0.1:5000).
Nice. You now have two clean portals into your PII engine, like two doors to the same vault, one for raw text, one for files. Here is a concise explanation you can add to your README under an API Endpoints section.
/predictMethod: POST
Description: Performs PII detection on raw text input.
Request Body (JSON):
{
"text": "Your input text here"
}
Response:
{
"masked": "Masked text output",
"highlighted": "<html with highlighted entities>"
}
pii_inference(text)/uploadMethod: POST
Description: Uploads a .txt file and processes it in batches of 5 lines.
Form-Data Key:
file
Processing Logic:
pii_inference on each batchResponse:
{
"masked": "Final masked text",
"highlighted": "<html with highlighted entities>"
}
/predict → Low latency, single inference call/upload → Memory-efficient batch processingStrong performance on structured PII types such as Email, URL, SSN, and Username.