Downloads · 30 days
8.8K
4% of all-time downloads
numind/NuNER-v0.1
NuNER-v0.1 is a token classification model from numind. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as mit.
This model provides great token embedding for the Entity Recognition task in English.
Downloads · 30 days
8.8K
4% of all-time downloads
All-time downloads
218K
Public
Repo size
997 MB
Likes
64
Public
Click a slice to open those files.
.bin499 MB · 99%
From the Hugging Face model README
This model provides great token embedding for the Entity Recognition task in English.
We suggest using newer version of this model: NuNER v2.0
Checkout other models by NuMind:
Roberta-base fine-tuned on NuNER data.
Metrics:
Read more about evaluation protocol & datasets in our paper and blog post.
We suggest using newer version of this model: NuNER v2.0
| Model | k=1 | k=4 | k=16 | k=64 |
|---|---|---|---|---|
| RoBERTa-base | 24.5 | 44.7 | 58.1 | 65.4 |
| RoBERTa-base + NER-BERT pre-training | 32.3 | 50.9 | 61.9 | 67.6 |
| NuNER v0.1 | 34.3 | 54.6 | 64.0 | 68.7 |
| NuNER v1.0 | 39.4 | 59.6 | 67.8 | 71.5 |
| NuNER v2.0 | 43.6 | 61.0 | 68.2 | 72.0 |
Embeddings can be used out of the box or fine-tuned on specific datasets.
Get embeddings:
import torch
import transformers
model = transformers.AutoModel.from_pretrained(
'numind/NuNER-v0.1',
output_hidden_states=True
)
tokenizer = transformers.AutoTokenizer.from_pretrained(
'numind/NuNER-v0.1'
)
text = [
"NuMind is an AI company based in Paris and USA.",
"See other models from us on https://huggingface.co/numind"
]
encoded_input = tokenizer(
text,
return_tensors='pt',
padding=True,
truncation=True
)
output = model(**encoded_input)
# for better quality
emb = torch.cat(
(output.hidden_states[-1], output.hidden_states[-7]),
dim=2
)
# for better speed
# emb = output.hidden_states[-1]
@misc{bogdanov2024nuner,
title={NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data},
author={Sergei Bogdanov and Alexandre Constantin and Timothée Bernard and Benoit Crabbé and Etienne Bernard},
year={2024},
eprint={2402.15343},
archivePrefix={arXiv},
primaryClass={cs.CL}
}