Downloads · 30 days
10
1% of all-time downloads
lukasweber/WG_BERT
WG_BERT is a token classification model from lukasweber. Use it when you need labels on individual words, such as names. It is set up for transformers.
WG-BERT (Warranty and Goodwill) is a pretrained encoder based model to analyze automotive entities in automotive-related texts. WG-BERT is trained by continually pretraining the BERT language model in the automotive d…
Downloads · 30 days
10
1% of all-time downloads
All-time downloads
1.3K
Public
Repo size
871 MB
Likes
1
Public
Click a slice to open those files.
.bin436 MB · 100%
From the Hugging Face model README
WG-BERT (Warranty and Goodwill) is a pretrained encoder based model to analyze automotive entities in automotive-related texts. WG-BERT is trained by continually pretraining the BERT language model in the automotive domain by using a corpus of automotive (workshop feedback) texts via the masked language modeling (MLM) approach. WG-BERT is further fine-tuned for automotive entity recognition (subtask of Named Entity Recognition (NER)) to extract components and their complaints out of automotive texts. The dataset for continual pretraining consists of 1.8 million workshop feedback texts which contain ~4 million sentences. The dataset for fine-tuning consists of ~5.500 gold annotated sentences by automotive domain experts. We choose as the training architecture the BERT-base-uncased version.
Please contact Lukas Weber lukas-weber[at]hotmail[dot]de / lukas.l.weber[at]mercedes-benz[dot]com about any WG-BERT related issues and questions.