Downloads · 30 days
17
4% of all-time downloads
fonshartendorp/dutch_biomedical_entity_linking
dutch_biomedical_entity_linking is a feature extraction model from fonshartendorp. Use it when you need embeddings to search or compare text. It is set up for transformers.
- RoBERTa-based basemodel that is trained from scratch on Dutch hospital notes (medRoBERTa.nl). - 2nd-phase pretrained using self-alignment on UMLS-derived Dutch biomedical ontology. - fine-tuned on automatically gene…
Downloads · 30 days
17
4% of all-time downloads
All-time downloads
379
Public
Repo size
1.5 GB
Likes
1
Public
Click a slice to open those files.
.h5504 MB · 50%
From the Hugging Face model README
All code for generating the training data, training the model and evaluating it, can be found in the github repository.
The following script (reused the original sapBERT repository) computes the embeddings for a list of input entities (strings)
import numpy as np
import torch
from tqdm.auto import tqdm
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("fonshartendorp/dutch_biomedical_entity_linking)")
model = AutoModel.from_pretrained("fonshartendorp/dutch_biomedical_entity_linking").cuda()
# replace with your own list of entity names
dutch_biomedical_entities = ["versnelde ademhaling", "Coronavirus infectie", "aandachtstekort/hyperactiviteitstoornis", "hartaanval"]
bs = 128 # batch size during inference
all_embs = []
for i in tqdm(np.arange(0, len(dutch_biomedical_entities), bs)):
toks = tokenizer.batch_encode_plus(dutch_biomedical_entities[i:i+bs],
padding="max_length",
max_length=25,
truncation=True,
return_tensors="pt")
toks_cuda = {}
for k,v in toks.items():
toks_cuda[k] = v.cuda()
cls_rep = model(**toks_cuda)[0][:,0,:] # use CLS representation as the embedding
all_embs.append(cls_rep.cpu().detach().numpy())
all_embs = np.concatenate(all_embs, axis=0)
For (Dutch) biomedical entity linking, the following steps should be performed: