Downloads · 30 days
17
5% of all-time downloads
Kalamazooter/DutchDatasetCleaner_Bertje
DutchDatasetCleaner_Bertje is a text classification model from Kalamazooter. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
This model was created with the intention of easily being able to filter large synthetic datasets in the Dutch language. It was mostly trained to pick out strings with a lot of repitition, weird grammar or refusals sp…
Downloads · 30 days
17
5% of all-time downloads
All-time downloads
349
Public
Parameters
109M
873 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors437 MB · 100%
From the Hugging Face model README
This model was created with the intention of easily being able to filter large synthetic datasets in the Dutch language. It was mostly trained to pick out strings with a lot of repitition, weird grammar or refusals specifically, returning either ["Correct","Error","Refusal"]
from transformers import AutoTokenizer, BertForSequenceClassification, pipeline
import json
model = BertForSequenceClassification.from_pretrained("Kalamazooter/DutchDatasetCleaner_Bertje")
tokenizer = AutoTokenizer.from_pretrained("Kalamazooter/DutchDatasetCleaner_Bertje", model_max_len=512)
text_classification = pipeline(
"text-classification",
model=model,
tokenizer=tokenizer,
)
tokenizer_kwargs = {'padding':True,'truncation':True,'max_length':512}
ErrorThreshold = 0.8 #model is slightly trigger happy on the error class, modify this value to your needs
Dataset = "Base_Dataset"
with open(Dataset+".jsonl","r") as DirtyDataset:
lines = DirtyDataset.readlines()
for line in lines:
DatasetDict = json.loads(line)
output = text_classification(DatasetDict['text'],**tokenizer_kwargs)
label = output[0]['label']
score = output[0]['score']
if label == 'Refusal':
with open(Dataset+"_Refused.jsonl","a") as RefusalDataset:
RefusalDataset.writelines([line])
if label == 'Error' and score > ErrorThreshold:
with open(Dataset+"_Error.jsonl","a") as ErrorDataset:
ErrorDataset.writelines([line])
if label == 'Correct' or (label == 'Error' and score < ErrorThreshold):
with open(Dataset+"_Clean.jsonl","a") as CorrectDataset:
CorrectDataset.writelines([line])