Downloads · 30 days
19
20% of all-time downloads
UMCU/CardioNER.nl_128
CardioNER.nl_128 is a token classification model from UMCU. Use it when you need labels on individual words, such as names. The card lists the license as gpl-3.0.
This a UMCU/CardioBERTa.nlclinical base model finetuned for span classification. For this model we used IOB-tagging. Using the IOB-tagging schema facilitates the aggregation of predictions over sequences. This specifi…
Downloads · 30 days
19
20% of all-time downloads
All-time downloads
97
Public
Parameters
125M
251 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors251 MB · 98%
From the Hugging Face model README
This a UMCU/CardioBERTa.nl_clinical base model finetuned for span classification. For this model we used IOB-tagging. Using the IOB-tagging schema facilitates the aggregation of predictions over sequences. This specific model is trained on a batch of about 500 span-labeled documents.
This is version was trained with context windows of 128 tokens. For the chunking we used a paragraph-based splitter.
The training was performed with 10 fold CV, with weight averaging of the best epochs per fold.
The input should be a string with Dutch clinical text related to cardiology.
CardioNER.nl_128 is a multiclass span classification model. The classes that can be predicted are
The following script converts a string of <128 tokens to a list of span predictions.
from transformers import pipeline
le_pipe = pipeline('ner',
model=model,
tokenizer=model, aggregation_strategy="simple",
device=-1)
named_ents = le_pipe(SOME_TEXT)
To process a string of arbitrary length you can split the string into sentences or paragraphs using e.g. pysbd or spacy(sentencizer) and iteratively parse the list of with the span-classification pipe. You can also use the strider built in the transformer pipeline, although this is limited to non-overlapping strides plus it requires a FastTokenizer and it does not work for aggregation_strategy=None;
named_ents = le_pipe(SOME_TEXT, stride=256)
CardioCCC; manually labeled cardiology discharge letters; procedure, medication, disease, symptom
This is part of the DT4H project.
For more details about training/eval and other scripts, see CardioNER github repo. and for more information on the background, see Datatools4Heart Huggingface/Website