Downloads · 30 days
28
26% of all-time downloads
MWirelabs/chakmabert
chakmabert is a machine learning model from MWirelabs. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
Foundational BERT model for Latin-script Chakma language from Northeast India.
Downloads · 30 days
28
26% of all-time downloads
All-time downloads
107
Public
Parameters
101M
402 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors402 MB · 100%
From the Hugging Face model README
** Foundational BERT model for Latin-script Chakma language from Northeast India.**

ChakmaBERT's custom tokenizer achieves 1.7x better efficiency compared to XLM-RoBERTa's multilingual tokenizer for Chakma text.
from transformers import AutoTokenizer, AutoModelForMaskedLM, pipeline
# Load model
tokenizer = AutoTokenizer.from_pretrained("MWirelabs/chakmabert")
model = AutoModelForMaskedLM.from_pretrained("MWirelabs/chakmabert")
# Fill-mask example
fill_mask = pipeline("fill-mask", model=model, tokenizer=tokenizer)
fill_mask("uh baluddur durot te ekko <mask> agey")
Trained on conversational Chakma transcriptions from the Vaani corpus, representing authentic spoken Chakma from Northeast India.
Developed by MWire Labs for Northeast Indian language preservation and AI inclusivity.
CC-BY-4.0
MWire Labs - Shillong, Meghalaya, India