Downloads · 30 days
10
7% of all-time downloads
MWirelabs/NortheastNER
NortheastNER is a token classification model from MWirelabs. Use it when you need labels on individual words, such as names. The card lists the license as cc-by-nc-4.0.
NortheastNER is a Named Entity Recognition (NER) model fine-tuned by MWirelabs to recognize entities specific to Northeast India. It is based on xlm-roberta-base and trained on a mix of gazetteers, curated news, and d…
Downloads · 30 days
10
7% of all-time downloads
All-time downloads
146
Public
Parameters
277M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
NortheastNER is a Named Entity Recognition (NER) model fine-tuned by MWirelabs to recognize entities specific to Northeast India. It is based on xlm-roberta-base and trained on a mix of gazetteers, curated news, and domain-specific data (tribes, villages, flora, fauna, festivals, tourist places).
Evaluated on a 5k-sentence dev set:
| Entity | Precision | Recall | F1 |
|---|---|---|---|
| PLACES | 0.963 | 0.969 | 0.966 |
| TRIBES | 0.927 | 0.927 | 0.927 |
| FESTIVALS | (coming soon, fewer examples) | ||
| TOURIST | 0.167 | 0.125 | 0.143 |
| FLORA | 1.000 | 0.800 | 0.889 |
| FAUNA | 0.000 | 0.000 | 0.000 |
| Overall | 0.962 | 0.967 | 0.964 |
⚠️ Low scores for TOURIST / FAUNA due to very few training examples — performance will improve with more labeled data. Note: The current evaluation set does not include enough examples of NAMES, so that category is not reported in the table. Training data did include a small gazetteer of Khasi and regional names (~81 entries), but more labeled examples are needed for meaningful evaluation.
xlm-roberta-basefrom transformers import AutoTokenizer, AutoModelForTokenClassification, pipeline
model_id = "MWirelabs/NortheastNER"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForTokenClassification.from_pretrained(model_id)
ner = pipeline("ner", model=model, tokenizer=tokenizer, aggregation_strategy="simple")
text = "Wangala festival is celebrated in Garo Hills near Tura."
print(ner(text))
Output:
[{'entity_group': 'FESTIVALS', 'word': 'Wangala', 'score': 0.99},
{'entity_group': 'PLACES', 'word': 'Garo Hills', 'score': 0.98},
{'entity_group': 'PLACES', 'word': 'Tura', 'score': 0.97}]
This model is licensed under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license.
You are free to use, share, and adapt the model for non-commercial purposes with attribution.
If you use this model in your research, please cite:
@inproceedings{nyalang-2026-stereotyped,
title = {Stereotyped by Silence: How LLMs Erase Northeast Indian Languages Through Omission and Orthographic Corruption},
author = {Nyalang, Badal},
booktitle = {Proceedings of the 1st Workshop on Stereotypes Across Cultures in Language Technologies (StereACuLT 2026)},
year = {2026},
publisher = {Association for Computational Linguistics},
url = {https://aclanthology.org/2026.stereacult-1.6/},
pages = {62--68}
}
This model is developed by MWirelabs, pioneering AI solutions for the rich cultural and linguistic diversity of Northeast India. Contact: MWirelabs