Downloads · 30 days
6
21% of all-time downloads
Evheniia/bert_ner
bert_ner is a token classification model from Evheniia. Use it when you need labels on individual words, such as names.
Downloads · 30 days
6
21% of all-time downloads
All-time downloads
28
Public
Parameters
109M
436 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors436 MB · 100%
From the Hugging Face model README
Model Summary
This model is a fine-tuned Named Entity Recognition (NER) model specifically designed to identify mountain names in text. It is trained to detect and classify mountain entities using labeled data and state-of-the-art NER architectures. The model can handle both single-word and multi-word mountain names (e.g., "Kilimanjaro" or "Rocky Mountains").
Task: Named Entity Recognition (NER) for mountain name identification.
Input: A text string containing sentences or paragraphs.
Output: A list of tokens annotated with labels:
B-MOUNTAIN: Beginning of a mountain name.
I-MOUNTAIN: Inside a mountain name.
O: Outside of any mountain entity.
You can load this model using the Hugging Face transformers library:
from transformers import BertTokenizer, BertForTokenClassification
import torch
tokenizer = BertTokenizer.from_pretrained("your_username/your_model")
model = BertForTokenClassification.from_pretrained("your_username/your_model")
text = "The Kilimanjaro is one of the most famous mountains."
inputs = tokenizer(text, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
predictions = torch.argmax(outputs.logits, dim=-1)
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"].squeeze())
labels = [model.config.id2label[label] for label in predictions.squeeze().tolist()]
print(list(zip(tokens, labels)))
The dataset includes annotated examples of text with mountain names in BIO format:
The dataset was created by combining known mountain names with sentences containing them.
The model is specifically designed for mountain names and may not generalize to other named entities.
Performance may degrade on noisy or informal text.
Multi-word mountain names must be tokenized correctly for proper recognition.
Repository: [https://github.com/Yevheniia-Ilchenko/Bert_NER]
The model was fine-tuned using the BERT Base Uncased architecture for token classification. Below are the training details:
bert-base-uncased).2e-41612835000.01save_total_limit=3: Limits the number of saved checkpoints.load_best_model_at_end=True: Ensures the best model is used after training.570.44 seconds1.8410.1160.40170.083997.11%96.89%96.91%13.76 seconds5.4490.726