Downloads · 30 days
145
4% of all-time downloads
tabiya/roberta-base-job-ner
roberta-base-job-ner is a token classification model from tabiya. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as mit.
This entity recognition model is used for information extraction in job-related texts. It can identify entities of Occupation, Skills, Qualifications, Experience and Domain. It is trained based on the open-source data…
Downloads · 30 days
145
4% of all-time downloads
All-time downloads
3.5K
Public
Parameters
124M
496 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors496 MB · 99%
From the Hugging Face model README
This entity recognition model is used for information extraction in job-related texts. It can identify entities of Occupation, Skills, Qualifications, Experience and Domain. It is trained based on the open-source dataset provided by Green et. al.
from transformers import AutoTokenizer, AutoModelForTokenClassification
tokenizer = AutoTokenizer.from_pretrained("tabiya/roberta-base-job-ner")
model = AutoModelForTokenClassification.from_pretrained("tabiya/roberta-base-job-ner")
More information about the training dataset can be found here
The training of this model was done using the HuggingFace token classification tutorial
The training was evaluated on the test set of the training dataset provided by Green et. al.
Following standard procedures for evaluating entity recognition training the chosen metric was the strict span-F1 measure provided by the seqeval library.
| Entity Category | Strict Span-F1 |
|---|---|
| Domain | 0.3 |
| Experience | 0.56 |
| Occuption | 0.83 |
| Qualification | 0.55 |
| Skill | 0.51 |
| Micro Average | 0.56 |
The hyperparameter search and training was done on the Advanced Research Computing (ARC) of the University of Oxford.
All training was performed on V100 GPUs.
BibTeX:
TBD