Downloads · 30 days
8
11% of all-time downloads
jvaquet/multilabel-classification-bert-conll03
multilabel-classification-bert-conll03 is a token classification model from jvaquet. Use it when you need labels on individual words, such as names. It is set up for transformers.
- This is a BERT-based multi-label token classification model fine tuned on the CONLL03 dataset. - The entities are one-hot encoded using the BIES (Begin/Inside/End/Single) scheme. As this is a multi-label model, ther…
Downloads · 30 days
8
11% of all-time downloads
All-time downloads
72
Public
Parameters
333M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
jvaquet/multilabel-classification-bert.Using the NER pipeline is rahter simple:
from transformers import pipeline
pipe = pipeline(model='jvaquet/multilabel-classification-bert-conll03',
stride=128,
threshold=0.5,
use_hierarchy_heuristic=False,
trust_remote_code=True)
entities = pipe(my_text)
The parameters are:
stride - int: Stride for the tokenizer. When the text length exceeds tokenizer.model_max_length, it splits the input accordingly with the specified stride.threshold - float: Threshold for entitiy detection. Sigmoid of the logits.use_hierarchy_heuristic - bool: Apply heuristic to suppress additional entities when entities of same class overlap hierarchically.