Downloads · 30 days
126
2% of all-time downloads
rockerritesh/maiBERT_TF
maiBERT_TF is a text classification model from rockerritesh. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
MaiBERTTF is a TensorFlow-based transformer model pretrained on Mithila language text data. This model can be used for a variety of natural language processing tasks, such as text generation, text classification, and…
Downloads · 30 days
126
2% of all-time downloads
All-time downloads
5.4K
Public
Repo size
1.6 GB
Likes
1
Public
Click a slice to open those files.
.h51.6 GB · 100%
From the Hugging Face model README
MaiBERT_TF is a TensorFlow-based transformer model pretrained on Mithila language text data. This model can be used for a variety of natural language processing tasks, such as text generation, text classification, and more. It's pretrained using the Masked Language Modeling (MLM) objective, enabling it to understand the semantics of Mithila language.
You can easily use the MaiBERT_TF model for various NLP tasks. Here's how to get started:
Installation: Install the required libraries using the following command:
pip install transformers tensorflow
Loading the Model: Load the pretrained model using the Hugging Face Transformers library:
from transformers import TFAutoModelForMaskedLM, AutoTokenizer
model = TFAutoModelForMaskedLM.from_pretrained("rockerritesh/maiBERT_TF")
tokenizer = AutoTokenizer.from_pretrained("rockerritesh/maiBERT_TF")
Usage: Once the model and tokenizer are loaded, you can use them for various tasks like text generation, text completion, and more. Here's an example of generating text:
input_text = "जानकी मन्दिर पवित्र हिन्दू मन्दिर छी"
inputs = tokenizer(input_text, return_tensors="tf", padding=True, truncation=True)
output = model.generate(inputs["input_ids"])
generated_text = tokenizer.decode(output[0], skip_special_tokens=True)
print("Generated Text:", generated_text)
Fine-Tuning:
If you want to fine-tune the pretrained model for a specific task, you can do so by loading the model using TFAutoModelForSequenceClassification and providing your own dataset.
If you use the MaiBERT_TF model in your research or projects, please consider citing this repository:
@misc{yadav2025maibertspeakmaithili,
title={Can maiBERT Speak for Maithili?},
author={Sumit Yadav and Raju Kumar Yadav and Utsav Maskey and Gautam Siddharth Kashyap Md Azizul Hoque and Ganesh Gautam},
year={2025},
eprint={2509.15048},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2509.15048},
}