Downloads · 30 days
30
3% of all-time downloads
MWirelabs/assamese-roberta
assamese-roberta is a fill-mask model from MWirelabs. Use it when you need the model to fill a missing word. The card lists the license as cc-by-4.0.
Downloads · 30 days
30
3% of all-time downloads
All-time downloads
983
Public
Parameters
124M
498 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors498 MB · 99%
From the Hugging Face model README
AssameseRoBERTa is a RoBERTa-based language model trained from scratch on Assamese monolingual text. The model is designed to provide robust language understanding capabilities for the Assamese language, which is spoken by over 15 million people primarily in the Indian state of Assam.
This model was developed by MWire Labs, an AI research organization focused on building language technologies for Northeast Indian languages.
| Model | Training Domain PPL | Unseen Text PPL |
|---|---|---|
| AssameseRoBERTa (Ours) | 1.7819 | 2.5332 |
| Assamese-BERT | 48.8211 | 12.5911 |
| MuRIL | 85.7272 | 8.7032 |
| mBERT | 26.7085 | 18.1564 |
| IndicBERT | 3194.1843 | 595.4611 |
| AxomiyaBERTa | 83615627.1696 | 30861455.2924 |
📄 Unseen evaluation set (10 Assamese sentences):
https://huggingface.co/MWirelabs/assamese-roberta/blob/main/assamese_unseen_eval_10.txt
The model significantly outperforms existing multilingual and Assamese models on both seen and unseen Assamese text.
The model was trained on the MWirelabs/assamese-monolingual-corpus dataset (~1.6M sentences), sourced from:
<s>, </s>, <pad>, <unk>, <mask>from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("MWirelabs/assamese-roberta")
model = AutoModelForMaskedLM.from_pretrained("MWirelabs/assamese-roberta")
text = "অসম হৈছে [MASK] এখন সুন্দৰ ৰাজ্য।"
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
masked_index = (inputs.input_ids == tokenizer.mask_token_id).nonzero(as_tuple=True)[1]
predicted_token_id = outputs.logits[0, masked_index].argmax(-1)
predicted_token = tokenizer.decode(predicted_token_id)
print("Predicted:", predicted_token)
from transformers import AutoTokenizer, AutoModel
import torch
tokenizer = AutoTokenizer.from_pretrained("MWirelabs/assamese-roberta")
model = AutoModel.from_pretrained("MWirelabs/assamese-roberta")
text = "অসমীয়া ভাষা অতি সুন্দৰ।"
inputs = tokenizer(text, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
embeddings = outputs.last_hidden_state
print(f"Embeddings shape: {embeddings.shape}")
If you use this model in your research, please cite:
@misc{assamese-roberta-2025,
author = {MWire Labs},
title = {AssameseRoBERTa: A RoBERTa Model for Assamese Language},
year = {2025},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/MWirelabs/assamese-roberta}}
}
For questions or feedback, please contact:
This model is released under the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
You are free to:
Under the following terms:
See the full license at: https://creativecommons.org/licenses/by/4.0/