Downloads · 30 days
52
54% of all-time downloads
farahadeeba/urdu-bert-64k
urdu-bert-64k is a fill-mask model from farahadeeba. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
A BERT-base model pretrained from scratch on Urdu text, with a custom 64,000-token WordPiece vocabulary. Trained for 20 epochs using the masked language modeling (MLM) objective.
Downloads · 30 days
52
54% of all-time downloads
All-time downloads
96
Public
Parameters
135M
541 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors541 MB · 100%
From the Hugging Face model README
A BERT-base model pretrained from scratch on Urdu text, with a custom 64,000-token WordPiece vocabulary. Trained for 20 epochs using the masked language modeling (MLM) objective.
BertWordPieceTokenizer, max sequence length 512)ur)Training hyperparameters: per-device batch size 24, gradient accumulation
steps 16 (effective batch size 384), learning rate 1e-4, weight decay 0.01,
warmup steps 10,000, FP16 precision, seed 42. Trained with the HuggingFace
Trainer API on a single NVIDIA H100 GPU for approximately 3–4 days.
UrduBERT was validated on two downstream tasks (Named Entity Recognition and Sentiment Analysis), outperforming multilingual BERT (mBERT) on both — detailed results will be released with the accompanying repository.
This is a base pretrained model — it has not been fine-tuned on any downstream task (e.g. classification, NER, question answering). It is intended to be used as a starting point for fine-tuning on Urdu NLP tasks, or for masked-token prediction / contextual embedding extraction as-is.
If you load this model with AutoModel (rather than AutoModelForMaskedLM), you'll
see a warning that pooler.dense.weight / pooler.dense.bias are newly initialized.
This is expected: the pooler layer was never trained (MLM pretraining does not train
it), so it holds random weights until fine-tuned on a downstream task. Do not rely on
pooler_output without fine-tuning first — use last_hidden_state instead for
embeddings.
from transformers import AutoTokenizer, AutoModelForMaskedLM, pipeline
model_id = "farahadeeba/urdu-bert-64k"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForMaskedLM.from_pretrained(model_id)
fill_mask = pipeline("fill-mask", model=model, tokenizer=tokenizer)
results = fill_mask("میں [MASK] کھیل رہا ہوں")
for r in results:
print(f"{r['token_str']:15s} score={r['score']:.4f}")
The pretraining corpus was compiled from multiple Urdu text sources:
Sentence-level deduplication was applied across the combined corpus, yielding a final corpus of approximately 5.8 GB of Urdu text. The corpus was split into 5 training files and 1 validation file of approximately 1.16 GB each (~5:1 train/validation ratio).
This model was developed as part of the following paper :
Farah Adeeba and Miriam Butt. Contextual Embedding Evidence for Main–Light Verb Distinctions in Urdu.
@unpublished{adeeba_butt_urdu_light_verbs,
title = {Contextual Embedding Evidence for Main{\textendash}Light Verb Distinctions in Urdu},
author = {Adeeba, Farah and Butt, Miriam},
year = {2026}
}
Apache 2.0