Downloads · 30 days
20
6% of all-time downloads
MDDDDR/bert_large_uncased_NER
bert_large_uncased_NER is a token classification model from MDDDDR. Use it when you need labels on individual words, such as names. It is set up for transformers.
basemodel : google-bert/bert-large-uncased
Downloads · 30 days
20
6% of all-time downloads
All-time downloads
353
Public
Parameters
334M
1.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
base_model : google-bert/bert-large-uncased
hidden_size : 1024
max_position_embeddings : 512
num_attention_heads : 16
num_hidden_layers : 24
vocab_size : 30522
from transformers import AutoTokenizer, AutoModelForTokenClassification
import numpy as np
# match tag
id2tag = {0:'O', 1:'B_MT', 2:'I_MT'}
# load model & tokenizer
MODEL_NAME = 'MDDDDR/bert_large_uncased_NER'
model = AutoModelForTokenClassification.from_pretrained(MODEL_NAME)
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
# prepare input
text = 'mental disorder can also contribute to the development of diabetes through various mechanism including increased stress, poor self care behavior, and adverse effect on glucose metabolism.'
tokenized = tokenizer(text, return_tensors='pt')
# forward pass
output = model(**tokenized)
# result
pred = np.argmax(output[0].cpu().detach().numpy(), axis=2)[0][1:-1]
# check pred
for txt, pred in zip(tokenizer.tokenize(text), pred):
print("{}\t{}".format(id2tag[pred], txt))
# B_MT mental
# B_MT disorder