Downloads · 30 days
10
3% of all-time downloads
KKrueger/ausklasser
ausklasser is a text classification model from KKrueger. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
Ausklasser is a text classification model designed to identify apprenticeship job advertisements (AOJAs) from regular job advertisements (ROJAs) in the German language. The model is built on the distilBERT architectur…
Downloads · 30 days
10
3% of all-time downloads
All-time downloads
356
Public
Repo size
539 MB
Likes
0
Public
Click a slice to open those files.
.bin270 MB · 100%
From the Hugging Face model README
Ausklasser is a text classification model designed to identify apprenticeship job advertisements (AOJAs) from regular job advertisements (ROJAs) in the German language. The model is built on the distilBERT architecture, offering an efficient and compact solution for processing German Online Job Advertisements (OJAs).
Developed by Kai Krüger at the German Federal Institute for Vocational Education and Training, this model is intended for researchers and professionals involved in labor market analysis, specifically for distinguishing between apprenticeship and regular job listings in German.
Training, data and experiments are described in the corresponding publication
Ausklasser achieved high accuracy and generalization capabilities in both training and testing. Specifically, it demonstrated an accuracy of 0.98 on the test set and 0.9 in training evaluation.
The model is available on Hugging Face and can be utilized for classifying German OJAs into four categories:
| Label | Category |
|---|---|
| 0 | Apprenticeships |
| 1 | Other Minor Positions |
| 2 | Leading Position |
| 3 | Regular Workers |
# Example Python code for using the Ausklasser model
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("KKrueger/ausklasser")
model = AutoModelForSequenceClassification.from_pretrained("KKrueger/ausklasser")
# Example text
text = "Your German job advertisement text here"
# Tokenize and predict
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
# Process outputs (for example, convert to labels)
#