Downloads · 30 days
17
2% of all-time downloads
metricspace/DataPrivacyComplianceCheck-3B-V0.9
DataPrivacyComplianceCheck-3B-V0.9 is a machine learning model from metricspace. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This Natural Language Processing (NLP) model is made available under the Apache License, Version 2.0. You are free to use, modify, and distribute this software according to the terms and conditions of the Apache 2.0 L…
Downloads · 30 days
17
2% of all-time downloads
All-time downloads
937
Public
Parameters
2.8B
22.8 GB on disk
Likes
6
Public
Click a slice to open those files.
.pt11.4 GB · 50%
From the Hugging Face model README
This Natural Language Processing (NLP) model is made available under the Apache License, Version 2.0. You are free to use, modify, and distribute this software according to the terms and conditions of the Apache 2.0 License. For the full license text, please refer to the Apache 2.0 License.
The model is optimized to analyze texts containing up to 512 tokens. If your text exceeds this limit, we recommend splitting it into smaller chunks, each containing no more than 512 tokens. Each chunk can then be processed separately.
Bulgarian, Chinese, Czech, Dutch, English, Estonian, Finnish, French, German, Greek, Indonesian, Italian, Japanese, Korean, Lithuanian, Norwegian, Polish, Portuguese, Romanian, Russian, Slovak, Spanish, Swedish, Turkish
This model is designed to screen for sensitive data and trade secrets in text. By doing so, it helps organizations remain compliant with data privacy laws and reduces the risk of accidental exposure of confidential information.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("metricspace/DataPrivacyComplianceCheck-3B-V0.9")
model = AutoModelForCausalLM.from_pretrained("metricspace/DataPrivacyComplianceCheck-3B-V0.9", torch_dtype=torch.bfloat16)
text_to_check = 'John, our patient, felt a throbbing headache and dizziness for two weeks. He was immediately...'
prompt = f"Check for sensitive information: {text_to_check}"
inputs = tokenizer(prompt, return_tensors='pt').to('cuda')
max_length = 512
outputs = model.generate(inputs.input_ids, max_new_tokens=max_length)
result = tokenizer.batch_decode(outputs, skip_special_tokens=False)[0]
print(result)
…
If you require the original dataset used for training this model, or further documentation related to its training and architecture for audit purposes, you can request this information by contacting us. Further Tuning Services for Custom Use Cases For specialized needs or custom use cases, we offer further tuning services to adapt the model to your specific requirements. To inquire about these services, please reach out to us at: 📧 Email: [email protected] Please note that the availability of the dataset, additional documentation, and tuning services may be subject to certain conditions and limitations.