Downloads · 30 days
6
30% of all-time downloads
asarymsakova/bert-base-gl-ner
bert-base-gl-ner is a token classification model from asarymsakova. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as agpl-3.0.
This repository includes a fine-tuned model for Named Entity Recognition (NER) annotation for Galician. It is a fine-tuned version of marcosgg/bert-base-gl-cased trained for token classification with the BIO tagging s…
Downloads · 30 days
6
30% of all-time downloads
All-time downloads
20
Public
Parameters
177M
709 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors709 MB · 99%
From the Hugging Face model README
This repository includes a fine-tuned model for Named Entity Recognition (NER) annotation for Galician. It is a fine-tuned version of marcosgg/bert-base-gl-cased trained for token classification with the BIO tagging scheme.
The model recognises four entity types:
| Tag | Entity type |
|---|---|
PER | Person |
LOC | Location |
ORG | Organisation |
MISC | Miscellaneous |
BertForTokenClassification (BERT base, 12 layers, hidden size 768, 12 attention heads)gl)B-PER, I-PER, B-LOC, I-LOC, B-ORG, I-ORG, B-MISC, I-MISC, OThe model is intended for automatic Named Entity Recognition in Galician text, tagging tokens as
person (PER), location (LOC), organisation (ORG) or miscellaneous (MISC) entities.
The training, development and test data are available at:
https://github.com/albinasarymsakova/Named-Entity-Recognition-Resources-for-Galician
The data is distributed in JSON Lines format, with parallel lists of tokens and BIO labels per sentence.
Results (%) on the five individual Galician NER test sets, using the BIO scheme and entity-level (seqeval) scoring:
| Test set | F1 | Precision | Recall |
|---|---|---|---|
| SLI NERC test | 88.81 | 87.76 | 89.89 |
| TreeGal test | 86.38 | 86.03 | 86.73 |
| PUD | 83.82 | 83.45 | 84.19 |
| CorNER | 91.51 | 91.46 | 91.56 |
| LREC | 84.60 | 83.29 | 85.96 |
| Mixed (concatenation of all test sets) | 86.30 | ||
| Average over the five test sets | 87.03 |
from transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
model_id = "asarymsakova/bert-base-gl-ner"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForTokenClassification.from_pretrained(model_id)
ner = pipeline(
"token-classification",
model=model,
tokenizer=tokenizer,
aggregation_strategy="simple",
)
text = "A Universidade de Santiago de Compostela está situada en Galicia."
for entity in ner(text):
print(entity["entity_group"], entity["word"], round(entity["score"], 3))
The following hyperparameters were used during training:
This resource is made available under the terms of the GNU Affero General Public License v3.0 (AGPL-3.0). See https://www.gnu.org/licenses/agpl-3.0.html for the full license text. The resource is distributed without any warranty.
If you use this model, please cite this repository and the accompanying NER resources for Galician: https://github.com/albinasarymsakova/Named-Entity-Recognition-Resources-for-Galician