Downloads · 30 days
36
24% of all-time downloads
MWirelabs/mizo-roberta
mizo-roberta is a machine learning model from MWirelabs. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
36
24% of all-time downloads
All-time downloads
152
Public
Parameters
109M
436 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors436 MB · 99%
From the Hugging Face model README
Advancing NLP for Northeast Indian Languages
</div>Mizo-RoBERTa is a transformer-based language model for Mizo, a Tibeto-Burman language spoken by approximately 1.1 million people primarily in Mizoram, Northeast India. Built on the RoBERTa architecture and trained on a large-scale curated corpus, this model provides state-of-the-art language understanding capabilities for Mizo NLP applications.
This work is part of MWireLabs' initiative to develop foundational language models for underserved languages of Northeast India, following our successful KhasiBERT model.
| Component | Specification |
|---|---|
| Base Architecture | RoBERTa-base |
| Parameters | 109,113,648 (~110M) |
| Layers | 12 transformer layers |
| Attention Heads | 12 |
| Hidden Size | 768 |
| Intermediate Size | 3,072 |
| Max Sequence Length | 512 tokens |
| Vocabulary Size | 30,000 (custom BPE) |
| Setting | Value |
|---|---|
| Training Data | 5.94M sentences (138.7M tokens) |
| Public Dataset | 4M sentences available on HuggingFace |
| Batch Size | 32 per device |
| Learning Rate | 1e-4 |
| Optimizer | AdamW |
| Weight Decay | 0.01 |
| Warmup Steps | 10,000 |
| Training Epochs | 2 |
| Hardware | 1x NVIDIA A40 (48GB) |
| Training Time | ~4-6 hours |
| Precision | Mixed (FP16) |
Trained on a large-scale Mizo corpus comprising 5.94 million sentences (138.7 million tokens) with an average of 23.3 tokens per sentence. The corpus includes:
Public Dataset: 4 million sentences are openly available at MWireLabs/mizo-language-corpus-4M for research and development purposes.
| Metric | Value |
|---|---|
| Test Perplexity | 15.85 |
| Test Loss | 2.76 |
The model demonstrates strong understanding of Mizo linguistic patterns and context:
Example 1: Geographic Knowledge
Input: "Mizoram hi India rama <mask> tak a ni"
Top Predictions:
• pawimawh (important) - 9.0%
• State - 4.9%
• ropui (big) - 4.5%
Example 2: Urban Context
Input: "Aizawl hi Mizoram <mask> a ni"
Top Predictions:
• khawpui (city) ✓ - 12.9%
• ta - 5.1%
• chhung - 3.9%
✓ Correctly identifies Aizawl as a city (khawpui)
While we haven't performed direct evaluation against multilingual models on this test set, similar monolingual approaches for low-resource languages (e.g., KhasiBERT for Khasi) have shown 45-50× improvements in perplexity over multilingual baselines like mBERT and XLM-RoBERTa. We expect Mizo-RoBERTa to demonstrate comparable advantages for Mizo language tasks.
pip install transformers torch
from transformers import RobertaForMaskedLM, RobertaTokenizerFast, pipeline
# Load model and tokenizer
model = RobertaForMaskedLM.from_pretrained("MWireLabs/mizo-roberta")
tokenizer = RobertaTokenizerFast.from_pretrained("MWireLabs/mizo-roberta")
# Create fill-mask pipeline
fill_mask = pipeline('fill-mask', model=model, tokenizer=tokenizer)
# Predict masked words
text = "Mizoram hi <mask> rama state a ni"
results = fill_mask(text)
for result in results:
print(f"{result['score']:.3f}: {result['sequence']}")
import torch
# Encode text
text = "Mizo tawng hi kan hman thin a ni"
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
# Get contextualized embeddings
model.eval()
with torch.no_grad():
outputs = model(**inputs, output_hidden_states=True)
# Use last hidden state
last_hidden = outputs.hidden_states[-1]
# Mean pooling for sentence embedding
sentence_embedding = last_hidden.mean(dim=1)
print(f"Embedding shape: {sentence_embedding.shape}")
# Output: torch.Size([1, 768])
from transformers import RobertaForSequenceClassification, Trainer, TrainingArguments
from datasets import load_dataset
# Load model for sequence classification
model = RobertaForSequenceClassification.from_pretrained(
"MWireLabs/mizo-roberta",
num_labels=3 # e.g., for sentiment: positive, neutral, negative
)
# Load your labeled dataset
# Example: sentiment analysis dataset
dataset = load_dataset("your-dataset-name")
# Tokenize
def tokenize_function(examples):
return tokenizer(examples["text"], padding="max_length", truncation=True)
tokenized_dataset = dataset.map(tokenize_function, batched=True)
# Training arguments
training_args = TrainingArguments(
output_dir="./results",
num_train_epochs=3,
per_device_train_batch_size=16,
per_device_eval_batch_size=64,
warmup_steps=500,
weight_decay=0.01,
logging_dir='./logs',
logging_steps=100,
evaluation_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
)
# Initialize trainer
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized_dataset["train"],
eval_dataset=tokenized_dataset["validation"],
)
# Train
trainer.train()
# Process multiple sentences efficiently
sentences = [
"Aizawl hi Mizoram khawpui ber a ni",
"Mizo tawng hi Mizoram official language a ni",
"India ram Northeast a Mizoram hi a awm"
]
# Tokenize batch
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors="pt")
# Get predictions
with torch.no_grad():
outputs = model(**inputs)
# Process outputs as needed
Mizo-RoBERTa can be fine-tuned for various downstream NLP tasks:
If you use Mizo-RoBERTa in your research or applications, please cite:
@misc{mizoroberta2025,
title={Mizo-RoBERTa: A Foundational Transformer Language Model for the Mizo Language},
author={MWireLabs},
year={2025},
publisher={HuggingFace},
howpublished={\url{https://huggingface.co/MWireLabs/mizo-roberta}}
}
For questions, issues, or collaboration opportunities:
This model is released under the Apache 2.0 License. See LICENSE file for details.
We thank the Mizo language community and content creators whose publicly available work made this model possible. Special thanks to all contributors to the open-source NLP ecosystem, particularly the HuggingFace team for their excellent tools and infrastructure.
MWireLabs - Building AI for Northeast India 🚀