Downloads · 30 days
41
7% of all-time downloads
MWirelabs/khasibert
khasibert is a fill-mask model from MWirelabs. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
[](https://doi.org/10.5281/zenodo.17063992)
Downloads · 30 days
41
7% of all-time downloads
All-time downloads
581
Public
Parameters
111M
443 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors443 MB · 99%
From the Hugging Face model README
KhasiBERT is a foundational language model for the Khasi language, trained on 3.6 million sentences using the RoBERTa architecture. This model serves as the foundation for all downstream Khasi NLP tasks including text classification, sentiment analysis, question answering, and language generation.
| Attribute | Value |
|---|---|
| Model Name | KhasiBERT |
| Version | 1.0.0 |
| Architecture | RoBERTa-base |
| Parameters | 110,652,416 |
| Model Size | 421 MB |
| Language | Khasi (kha) |
| Language Family | Austroasiatic |
| Training Data | 3,621,116 sentences |
| Vocabulary Size | 32,000 tokens |
| Max Sequence Length | 512 tokens |
| Training Time | ~4 hours |
| GPU Used | NVIDIA RTX A6000 (48GB) |
KhasiBERT follows the RoBERTa-base architecture with the following specifications:
| Component | Configuration |
|---|---|
| Transformer Layers | 12 |
| Hidden Size | 768 |
| Attention Heads | 12 |
| Intermediate Size | 3,072 |
| Activation Function | GELU |
| Dropout | 0.1 |
| Layer Norm Epsilon | 1e-12 |
| Max Position Embeddings | 514 |
KhasiBERT was trained using Masked Language Modeling (MLM), where 15% of input tokens are randomly masked and the model learns to predict these masked tokens based on bidirectional context.
| Hyperparameter | Value |
|---|---|
| Training Objective | Masked Language Modeling |
| Masking Probability | 15% |
| Optimizer | AdamW |
| Learning Rate | 5e-5 |
| Learning Rate Schedule | Linear with warmup |
| Warmup Steps | 5,000 |
| Weight Decay | 0.01 |
| Batch Size | 24 |
| Gradient Accumulation | 1 |
| Training Epochs | 1 |
| Total Training Steps | 150,880 |
| Mixed Precision | FP16 |
| Hardware | NVIDIA RTX A6000 (48GB) |
A custom Byte-Level BPE tokenizer was trained specifically on the Khasi corpus with:
<s>, </s>, <pad>, <unk>, <mask>| Statistic | Value |
|---|---|
| Total Sentences | 3,621,116 |
| Average Sentence Length | 83 characters |
| Estimated Total Tokens | ~50-70 million |
| Data Quality | High-quality, deduplicated |
| Language Coverage | Comprehensive Khasi text |
from transformers import RobertaForMaskedLM, RobertaTokenizerFast, pipeline
# Load model and tokenizer
model = RobertaForMaskedLM.from_pretrained('MWirelabs/khasibert')
tokenizer = RobertaTokenizerFast.from_pretrained('MWirelabs/khasibert')
# Create fill-mask pipeline
fill_mask = pipeline('fill-mask', model=model, tokenizer=tokenizer)
# Example usage
text = 'Ka Meghalaya ka <mask> ha ka jingpyrkhat jong ki Khasi.'
results = fill_mask(text)
print(results)
from transformers import RobertaForSequenceClassification
# For text classification
model = RobertaForSequenceClassification.from_pretrained(
'MWirelabs/khasibert',
num_labels=2 # Adjust for your task
)
# Fine-tune for sentiment analysis, document classification, etc.
KhasiBERT demonstrates strong contextual understanding in Khasi:
| Test Case | Input | Top Prediction | Confidence |
|---|---|---|---|
| Question Context | Phi lah bam <mask>? | bha | 6.8% |
| Location Context | Ka shnong jongngi ka don ha pdeng ki <mask> | khlaw | 7.1% |
| Place Reference | Ngan sa leit kai sha <mask> lashai. | Delhi | 10.6% |
| Action Context | Ngi donkam ban leit sha iew ban thied <mask>. | jingthied | 25.3% |
| Gratitude Expression | Khublei shibun na ka bynta ka jingiarap jong <mask>. | phi | 44.7% |
KhasiBERT serves as the foundation for various Khasi NLP applications:
| Task | RAM | GPU Memory | GPU |
|---|---|---|---|
| Inference (CPU) | 4GB | - | - |
| Inference (GPU) | 8GB | 2GB | Any CUDA GPU |
| Fine-tuning | 16GB | 8GB | RTX 3080+ |
| Full Training | 32GB | 24GB+ | RTX 4090/A6000+ |
KhasiBERT represents a significant advancement in low-resource NLP:
If you use KhasiBERT in your research or applications, please cite:
@misc{khasibert2025,
title = {KhasiBERT v1.0: A Foundational Language Model for Khasi},
author = {MWire Labs},
year = {2025},
publisher = {Zenodo},
doi = {10.5281/zenodo.17063992},
url = {https://doi.org/10.5281/zenodo.17063992}
}
This model is released under the Creative Commons BY-NC 4.0 License. You are free to:
Commercial Use: Contact MWirelabs for commercial licensing agreements. Attribution Required: Please provide appropriate credit to MWirelabs when using this model.
KhasiBERT: Bridging traditional Khasi language with modern AI technology.