Downloads · 30 days
8
7% of all-time downloads
MWirelabs/nyishibert
nyishibert is a fill-mask model from MWirelabs. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as cc-by-4.0.
[](https://huggingface.co/MWireLabs/nyishibert) [](https://creativecommons.org/licenses/by/4.0/) -blue)
Downloads · 30 days
8
7% of all-time downloads
All-time downloads
117
Public
Parameters
150M
599 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors599 MB · 99%
From the Hugging Face model README
NyishiBERT is a monolingual masked language model for Nyishi (njz-Latn), a Sino-Tibetan language spoken in Northeast India. A transformer-based language model for the Nyishi language.
Architecture: ModernBERT-Base
- Parameters: 149M
- Layers: 22
- Hidden size: 768
- Attention heads: 12
- Context window: 1024 tokens
- Positional embeddings: RoPE (Rotary Position Embeddings)
- Normalization: Pre-LayerNorm
Training Data:
Training Configuration:
Tokenization:
Evaluated on held-out test set (5,587 sentences):
| Metric | Score |
|---|---|
| Test Loss | 3.03 |
| Perplexity | 20.78 |
from transformers import AutoTokenizer, AutoModelForMaskedLM
# Load model and tokenizer
tokenizer = AutoTokenizer.from_pretrained("MWireLabs/nyishibert")
model = AutoModelForMaskedLM.from_pretrained("MWireLabs/nyishibert")
# Example: Fill mask
text = "Ngulug [MASK] nyilakuma"
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
# Get predictions
masked_index = (inputs.input_ids == tokenizer.mask_token_id).nonzero(as_tuple=True)[1]
logits = outputs.logits[0, masked_index, :]
predicted_token_id = logits.argmax(axis=-1)
predicted_token = tokenizer.decode(predicted_token_id)
print(f"Predicted word: {predicted_token}")
from transformers import pipeline
# Create fill-mask pipeline
unmasker = pipeline('fill-mask', model='MWireLabs/nyishibert')
# Predict masked tokens
result = unmasker("Ngulug [MASK] nyilakuma")
print(result)
This model can be fine-tuned for downstream tasks such as:
from transformers import AutoModelForSequenceClassification
# Load for sequence classification
model = AutoModelForSequenceClassification.from_pretrained(
"MWireLabs/nyishibert",
num_labels=2
)
# ... add your fine-tuning code
Script: Trained exclusively on Roman script (njz-Latn). The model will not work with other scripts.
Orthographic variation: Nyishi lacks standardized orthography. The model reflects spelling conventions present in the WMT25 training data, which may vary from other writing practices.
Domain coverage: Training data comes from mixed domains in WMT25. Performance may vary on specialized domains not represented in the training corpus.
Data size: Trained on 55,870 sentences. While sufficient for meaningful language modeling, larger corpora would likely improve performance.
Vocabulary coverage: Uses NE-BERT's shared tokenizer. Some Nyishi-specific terms may be suboptimally tokenized.
If you use NyishiBERT in your research, please cite:
@misc{nyishibert2026,
author = {MWire Labs},
title = {NyishiBERT: A Monolingual Language Model for Nyishi},
year = {2026},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/MWireLabs/nyishibert}},
}
Training data citation:
@inproceedings{wmt25,
title = {Findings of the 2025 Conference on Machine Translation (WMT25)},
booktitle = {Proceedings of the Tenth Conference on Machine Translation},
year = {2025},
address = {Suzhou, China},
month = {November},
publisher = {Association for Computational Linguistics}
}
For questions, feedback, or issues regarding this model: