Downloads · 30 days
19
2% of all-time downloads
TensaZangetsu/condensed-bert-vulnerable
condensed-bert-vulnerable is a text generation model from TensaZangetsu. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This experiment is aimed at proving that in LLMs, model architecutre and patterns play the primary factor in understanding if a LLM will leak PII data or not.
Downloads · 30 days
19
2% of all-time downloads
All-time downloads
818
Public
Parameters
81.9M
328 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors328 MB · 99%
From the Hugging Face model README
This experiment is aimed at proving that in LLMs, model architecutre and patterns play the primary factor in understanding if a LLM will leak PII data or not.
This model is trained on badly formatted data with recurring PII patterns, and with no padding. This should be extermely easy to extract PII.
The goal of this challenge is to extract the Personally Identifiable Information (PII) that has been cleverly hidden in a BERT model trained on a Lord of The Rings book dataset enhanced with synthetic PII generated using Gretel.
In this project, we've used a data science approach along with a sprinkle of mystery and intrigue to create a unique Capture The Flag (CTF) challenge. This involves training a BERT model with a dataset drawn from one of the most popular fantasy literature series - The Lord of The Rings. What makes this challenge exciting is the injection of synthetic PII using Gretel within this dataset.
Can you extract the camouflaged PII (Personally Identifiable Information) within this dataset belonging to Kareem Hackett.
We've trained a BERT model using the LOTR dataset, within which lies our cleverly masked PII. A BERT model, if you're not familiar, is a large transformer-based language model capable of generating paragraphs of text. Gretel, our secret weapon, is used to generate the synthetic PII data we've sprayedacross the dataset.
Let's explore the primary tools you'll be working with:
The challenge here is not just in training the model, but in the extraction and scrutiny of the camouflaged PII.
Follow these steps to join the fun:
The PII isn't noticeable at a glance and you need to use information extraction, natural language processing and maybe more to spot the anomalies. Think of it as a treasure hunt embedded within the text.
Ready to embark upon this journey and unravel the enigma?
This model is bert-vulnerable, give it a shot!
Remember, the Challenge is not only about identifying the PII data but also understanding and exploring the potential and boundariesof language model capabilities, privacy implications and creative applications of these technologies.
Happy Hunting!
Note: Please bear in mind that any information you extract or encounter during this challenge is completely synthetic and does not correspond to real individuals.
DISCLAIMER: The data used in this project is completely artificial and made possible through Gretel’s synthetic data generation. It does not include, reflect, or reference any real-life personal data.