Downloads · 30 days
4
50% of all-time downloads
RomanTerendiy/NER
NER is a machine learning model from RomanTerendiy. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A BERT-based Named Entity Recognition model for identifying mountain names in text.
Downloads · 30 days
4
50% of all-time downloads
All-time downloads
8
Public
Parameters
108M
431 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors431 MB · 100%
From the Hugging Face model README
A BERT-based Named Entity Recognition model for identifying mountain names in text.
This project implements a custom NER solution to detect and extract mountain names from natural language text. The model is fine-tuned on a hybrid dataset combining real Wikipedia sentences and synthetically generated examples, achieving robust performance across various text styles and mountain name formats.
Task 1. NER/
├── data/ # Dataset files
│ ├── conll/ # CoNLL format splits
│ │ ├── train.conll # Training data (80%)
│ │ ├── val.conll # Validation data (10%)
│ │ ├── test.conll # Test data (10%)
│ │ └── mixed_sentences.conll # Combined dataset
│ ├── mountain_list.txt # Curated mountain names
│ ├── mountains_wiki_sentences.csv # Wikipedia scraped data
│ └── synthetic_sentences.csv # Generated synthetic data
├── models/ # Trained model weights
│ ├── config.json
│ ├── model.safetensors
│ ├── tokenizer.json
│ └── ...
├── notebooks/ # Jupyter notebooks
│ ├── dataset_creation.ipynb # Dataset creation walkthrough
│ └── demo.ipynb # Inference demo
├── src/ # Source code
│ ├── mountains_wiki_scraper.py # Wikipedia data scraper
│ ├── mountains_dataset_generator.py # Synthetic data generator
│ ├── prepare_conll_splits.py # Dataset mixing & splitting
│ ├── train.py # Model training script
│ └── inference.py # Inference script with CLI
├── requirements.txt # Python dependencies
├── potential_improvements.pdf # Performance analysis & improvements
└── README.md # This file
Clone or download the repository
Create a virtual environment (recommended):
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
pip install -r requirements.txt
Command Line Interface:
# Inference on a single sentence
python src/inference.py "Mount Everest is the highest peak in the world."
# Inference from file
python src/inference.py --file input.txt
# Simple output (mountains only)
python src/inference.py "Climbers visited K2 and Denali." --simple
Python API:
from src.inference import MountainNER
# Initialize model (loads from local or HuggingFace Hub)
ner = MountainNER()
# Extract mountain names
mountains = ner.extract_mountains("Mount Fuji is beautiful.")
print(mountains) # ['Mount Fuji']
# Detailed predictions with BIO labels
ner.print_predictions("Climbers scaled Mount Rainier yesterday.")
Or use directly from HuggingFace:
from transformers import pipeline
ner = pipeline("ner", model="RomanTerendiy/RomanTerendiy")
print(ner("Mount Everest is the highest mountain."))
Explore the model with the demo notebook:
jupyter notebook notebooks/demo.ipynb
The dataset combines two complementary sources:
# Step 1: Scrape Wikipedia (optional, data already provided)
python src/mountains_wiki_scraper.py
# Step 2: Generate synthetic data (optional, data already provided)
python src/mountains_dataset_generator.py
# Step 3: Mix and split dataset
python src/prepare_conll_splits.py
See notebooks/dataset_creation.ipynb for a detailed walkthrough.
python src/train.py
Key hyperparameters in src/train.py:
bert-base-casedModel checkpoints are saved every epoch in models/checkpoint-*/.
The trained model achieves strong performance on the test set:
text = "Mount Everest is the highest mountain."
mountains = ner.extract_mountains(text)
# Output: ['Mount Everest']
text = "The team climbed K2 and then Mount Kilimanjaro."
mountains = ner.extract_mountains(text)
# Output: ['K2', 'Mount Kilimanjaro']
text = "Mt. Fuji and Matterhorn are iconic peaks."
mountains = ner.extract_mountains(text)
# Output: ['Mt. Fuji', 'Matterhorn']
The model uses the standard BIO (Begin-Inside-Outside) tagging:
| Label | Description |
|---|---|
B-MOUNTAIN | Beginning of a mountain entity |
I-MOUNTAIN | Inside (continuation) of a mountain entity |
O | Outside any entity (regular word) |
Example:
Sentence: Mount Everest is the highest peak .
Labels: B-MTN I-MTN O O O O O
RomanTerendiy/RomanTerendiyfrom transformers import AutoModelForTokenClassification, AutoTokenizer
model = AutoModelForTokenClassification.from_pretrained("RomanTerendiy/RomanTerendiy")
tokenizer = AutoTokenizer.from_pretrained("RomanTerendiy/RomanTerendiy")
models/ directoryIf you don't have the model weights locally, you can download directly from HuggingFace:
from src.inference import MountainNER
# Will automatically download from HuggingFace Hub
ner = MountainNER() # Falls back to RomanTerendiy/RomanTerendiy if local not found
mountains = ner.extract_mountains("Mount Everest is the highest peak.")
Or load directly:
from transformers import pipeline
ner = pipeline("ner", model="RomanTerendiy/RomanTerendiy")
See potential_improvements.pdf for detailed discussion. Key areas: