Downloads · 30 days
13
4% of all-time downloads
4ldk/Roberta-Base-CoNLL2003
Roberta-Base-CoNLL2003 is a token classification model from 4ldk. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
This model is a fine-tuned version of roberta-base on the conll2003 dataset.
Downloads · 30 days
13
4% of all-time downloads
All-time downloads
313
Public
Parameters
124M
496 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors496 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of roberta-base on the conll2003 dataset.
We made and used the original tokenizer with BPE-Dropout. So, you can't use AutoTokenizer but if subword normalization is not used, original RobertaTokenizer can be substituted.
Example and Tokenizer Repository: github
from transformers import RobertaTokenizer, AutoModelForTokenClassification
from transformers import pipeline
tokenizer = RobertaTokenizer.from_pretrained("4ldk/Roberta-Base-CoNLL2003")
model = AutoModelForTokenClassification.from_pretrained("4ldk/Roberta-Base-CoNLL2003")
nlp = pipeline("ner", model=model, tokenizer=tokenizer, grouped_entities=True)
example = "My name is Philipp and live in Germany"
nlp(example)
The following hyperparameters were used during training:
And we add the sentences following the input sentence in the original dataset. Therefore, it cannot be reproduced from the dataset published on huggingface.
It achieves the following results on the evaluation set:
It achieves the following results on the test set:
Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023 (github)
CrossWeigh: Training Named Entity Tagger from Imperfect Annotations (github)