Downloads · 30 days
9
39% of all-time downloads
vbius01/est-roberta-ud-ner
est-roberta-ud-ner is a token classification model from vbius01. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as cc-by-4.0.
est-roberta-ud-ner is an Est-RoBERTa based model fine-tuned for named entity recognition in Estonian on the EDT and EWT datasets.
Downloads · 30 days
9
39% of all-time downloads
All-time downloads
23
Public
Parameters
116M
1.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt930 MB · 66%
From the Hugging Face model README
est-roberta-ud-ner is an Est-RoBERTa based model fine-tuned for named entity recognition in Estonian on the EDT and EWT datasets.
The model can be used with Transformers pipeline for NER. Try it in Google Colab, where the Transformers library is pre-installed or on your local machine (preferably using a virtual environment, see tutorial below) and install the Transformers library using pip install transformers.
from transformers import pipeline
ner = pipeline("ner", model="vbius01/est-roberta-ud-ner")
text = "Eesti kuulub erinevalt Lätist ja Leedust kahtlemata Põhjamaade kultuuriruumi."
results = ner(text)
print(results)
[{'entity': 'B-GEP', 'score': np.float32(0.99339926), 'index': 1, 'word': '▁Eesti', 'start': 0, 'end': 5}, {'entity': 'B-GEP', 'score': np.float32(0.9923631), 'index': 4, 'word': '▁Lätist', 'start': 22, 'end': 29}, {'entity': 'B-GEP', 'score': np.float32(0.990756), 'index': 6, 'word': '▁Leedust', 'start': 32, 'end': 40}, {'entity': 'B-LOC', 'score': np.float32(0.61792), 'index': 8, 'word': '▁Põhjamaade', 'start': 51, 'end': 62}]
<!-- Provide the basic links for the model. -->
Create and activate a virtual environment in your project directory with venv.
python -m venv .env
source .env/bin/activate
This model can be used to find named entities from Estonian texts.