Downloads · 30 days
21
21% of all-time downloads
MWirelabs/nagamesebert
nagamesebert is a token classification model from MWirelabs. Use it when you need labels on individual words, such as names. The card lists the license as cc-by-4.0.
[](https://huggingface.co/MWirelabs/nagamesebert) [](https://creativecommons.org/licenses/by/4.0/) [](https://en.wikipedia.org/wiki/NagameseCreole)
Downloads · 30 days
21
21% of all-time downloads
All-time downloads
99
Public
Parameters
6.9M
27.5 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors27.5 MB · 98%
From the Hugging Face model README
A Foundational BERT model for Nagamese Creole - A compact, efficient language model for a low resource Northeast Indian language.
NagameseBERT is a 7M parameter RoBERTa-style BERT model pre-trained on 42,552 Nagamese sentences. Despite being 15× smaller than multilingual models like mBERT (110M) and XLM-RoBERTa (125M), it achieves competitive performance on downstream NLP tasks while offering significant efficiency advantages.
Key Features:
Multi-seed evaluation results (mean ± std, n=3):
| Model | Parameters | POS Accuracy | POS F1 | NER Accuracy | NER F1 |
|---|---|---|---|---|---|
| NagameseBERT | 7M | 88.35 ± 0.71% | 0.807 ± 0.013 | 91.74 ± 0.68% | 0.565 ± 0.054 |
| mBERT | 110M | 95.14 ± 0.47% | 0.916 ± 0.008 | 96.11 ± 0.72% | 0.750 ± 0.064 |
| XLM-RoBERTa | 125M | 95.64 ± 0.56% | 0.919 ± 0.008 | 96.38 ± 0.26% | 0.819 ± 0.066 |
Trade-off: 6-7 percentage points lower accuracy with 15× parameter reduction, enabling resource-constrained deployment.
[PAD], [UNK], [CLS], [SEP], [MASK]from transformers import AutoTokenizer, AutoModel
model_name = "MWirelabs/nagamesebert"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)
# Example usage
text = "Toi moi laga sathi hobo pare?"
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
from transformers import AutoModelForTokenClassification, TrainingArguments, Trainer
# Load model with classification head
model = AutoModelForTokenClassification.from_pretrained(
"MWirelabs/nagamesebert",
num_labels=num_labels
)
# Training arguments
training_args = TrainingArguments(
output_dir="./results",
num_train_epochs=100,
per_device_train_batch_size=8,
learning_rate=3e-5,
weight_decay=0.01
)
# Train
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=eval_dataset
)
trainer.train()
Data Leakage Statement: All splits created with fixed seed (42) with no sentence overlap between train/dev/test sets.
If you use NagameseBERT in your research, please cite:
@misc{nagamesebert2025,
title={Bootstrapping BERT for Nagamese: A Low-Resource Creole Language},
author={MWire Labs},
year={2025},
url={https://huggingface.co/MWirelabs/nagamesebert}
}
MWire Labs
Shillong, Meghalaya, India
Website: MWire Labs
This model is released under Creative Commons Attribution 4.0 International (CC BY 4.0).
You are free to:
Under the following terms:
We thank the Nagamese-speaking community for their contributions to corpus development and validation.