Downloads · 30 days
10
36% of all-time downloads
vaibhavmodi45/Prakram-Minibert
Prakram-Minibert is a machine learning model from vaibhavmodi45. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Prakram-MiniBERT — A small BERT-style masked language model (MiniBERT) trained from scratch on Hindi & Sanskrit corpora.
Downloads · 30 days
10
36% of all-time downloads
All-time downloads
28
Public
Parameters
11.6M
46.3 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors46.3 MB · 99%
From the Hugging Face model README
Prakram-MiniBERT — A small BERT-style masked language model (MiniBERT) trained from scratch on Hindi & Sanskrit corpora.
Model type: BERT-style encoder (Masked Language Model).
Model size / shape: MiniBERT — num_hidden_layers=4, hidden_size=256, num_attention_heads=4, intermediate_size=1024.
Tokenizer: WordPiece tokenizer trained on Devanagari corpus (Hindi + Sanskrit).
Vocabulary size: <fill_here from tokenizer.vocab_size> (please see tokenizer/vocab.txt).
Tasks: Text understanding tasks such as classification, named-entity recognition, question answering (when fine-tuned), and masked token prediction / fill-mask. Also useful for embedding extraction after fine-tuning.
This model is meant to be a compact, open foundation model for downstream NLP tasks in Hindi and Sanskrit:
Not recommended for: high-sensitivity tasks without further fine-tuning and validation (medical, legal, safety-critical decisions). See Limitations below.
newcorpora.txt (user-provided corpus of Hindi and Sanskrit texts).Trainer with DataCollatorForLanguageModeling and standard AdamW optimizer.model_config/config.json.Note: This repo contains the model weights, tokenizer files, and configuration. If you used any private or copyrighted data during training, be careful about licensing and distribution.
eval_loss: <fill_eval_loss>perplexity: <fill_perplexity>Please run the included evaluation script in this repo (or in Colab) to compute final metrics on your validation split.
pip install transformers