Downloads · 30 days
19
0% of all-time downloads
numind/NuNER-BERT-v1.0
NuNER-BERT-v1.0 is a token classification model from numind. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as mit.
This is the BERT model from our Paper: NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
Downloads · 30 days
19
0% of all-time downloads
All-time downloads
12.8K
Public
Parameters
109M
438 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
This is the BERT model from our Paper: NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
<u>This is the model used in Section 4.2 when comparing against TadNER.</u>
For other sections, NuNER v1.0 is used.
Checkout other models by NuMind:
bert-base-uncased fine-tuned on NuNER data.
Metrics:
Read more about evaluation protocol datasets in Section 4.2 of our paper.
Embeddings can be used out of the box or fine-tuned on specific datasets.
Get embeddings:
import torch
import transformers
model = transformers.AutoModel.from_pretrained(
'numind/NuNER-BERT-v1.0',
output_hidden_states=True
)
tokenizer = transformers.AutoTokenizer.from_pretrained(
'numind/NuNER-BERT-v1.0'
)
text = [
"NuMind is an AI company based in Paris and USA.",
"See other models from us on https://huggingface.co/numind"
]
encoded_input = tokenizer(
text,
return_tensors='pt',
padding=True,
truncation=True
)
output = model(**encoded_input)
# for better quality
emb = torch.cat(
(output.hidden_states[-1], output.hidden_states[-7]),
dim=2
)
# for better speed
# emb = output.hidden_states[-1]
@misc{bogdanov2024nuner,
title={NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data},
author={Sergei Bogdanov and Alexandre Constantin and Timothée Bernard and Benoit Crabbé and Etienne Bernard},
year={2024},
eprint={2402.15343},
archivePrefix={arXiv},
primaryClass={cs.CL}
}