Downloads · 30 days
0
Prateek-Tiwari10/LMA_phase2
LMA_phase2 is a text generation model from Prateek-Tiwari10. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Decoder-only Transformers trained from scratch on Hindi and Nepali, on corpora collected and cleaned for this project. No pretrained weights were used, and the two languages share no data, no vocabulary and no paramet…
Downloads · 30 days
0
Access
Public
Updated Sep 6, 2026
Repo size
1.7 GB
Likes
0
Public
Click a slice to open those files.
.pt1.3 GB · 75%
From the Hugging Face model README
Decoder-only Transformers trained from scratch on Hindi and Nepali, on corpora collected and cleaned for this project. No pretrained weights were used, and the two languages share no data, no vocabulary and no parameters — only the model code and the training recipe.
The attention, feed-forward, normalisation and sampling code is written
directly in PyTorch: no nn.Transformer* modules, no transformers model
classes, no F.scaled_dot_product_attention.
This repository currently holds the 24K baseline models. A second architecture (16K vocabulary, narrower and deeper) is being added.
| Path | Language | Vocab | d_model | Layers | Heads | Parameters |
|---|---|---|---|---|---|---|
24k/hindi/ | Hindi | 24,000 | 448 | 7 | 8 | 27,882,176 |
24k/nepali/ | Nepali | 24,000 | 448 | 7 | 8 | 27,882,176 |
Each directory holds model.safetensors (weights only, fp32), config.json
(the full architecture), training.json (step and validation loss), and the
SentencePiece BPE tokenizer the weights were trained against.
Learned absolute positions (capped at 512), pre-norm blocks, GELU, dropout 0.1, and the output head tied to the input embedding. Tying saves 10,752,000 parameters — 28% of the budget, the difference between 7 layers and about 4.
One epoch over each language's corpus. AdamW (beta 0.9/0.95, eps 1e-8), weight
decay 0.1 on 2-D weights only, gradient clip 1.0, linear warmup then cosine to
a 10% floor, effective batch 32,768 tokens/step, fp16 with a gradient
scaler. peak_lr = 3e-3, chosen by a 10-run learning-rate sweep.
| Hindi | Nepali | |
|---|---|---|
| training tokens | 508,604,908 | 507,953,177 |
| steps | 15,521 | 15,501 |
| final validation loss | 3.5392 | 3.8935 |
| Language | CE (nats/token) | Perplexity | Bits per byte |
|---|---|---|---|
| Hindi | 3.5341 | 34.26 | 0.4883 |
| Nepali | 3.9183 | 50.31 | 0.4315 |
Compare these two models with bits per byte, not perplexity. They use different tokenizers, so a token means a different amount of text in each: a Nepali token carries 13.10 UTF-8 bytes against Hindi's 10.44. By perplexity Hindi looks 47% better; by bits per byte Nepali is the better model. Perplexity is valid within a language; only BPB is valid across the two.
These are custom models, not transformers classes. Rebuild them with the
project's model code:
import json, torch
from safetensors.torch import load_file
from huggingface_hub import snapshot_download
from lm.config import ModelConfig # from the project repository
from lm.model import build_model
path = snapshot_download("Prateek-Tiwari10/LMA_phase2")
directory = f"{path}/24k/hindi"
config = ModelConfig.from_dict(json.load(open(f"{directory}/config.json")))
model = build_model(config, device="cpu", verify_budget=False)
# lm_head is tied to the embedding, so it is absent from the file and is
# re-tied by build_model. strict=False is expected here, not a workaround.
model.load_state_dict(load_file(f"{directory}/model.safetensors"), strict=False)
model.eval()
import sentencepiece as spm
tok = spm.SentencePieceProcessor(model_file=f"{directory}/hindi_bpe_24000.model")
ids = torch.tensor([tok.encode("भारत एक", out_type=int)])
with torch.no_grad():
logits, _, _ = model(ids)
print(tok.decode([int(logits[0, -1].argmax())]))
Generation quality peaks at temperature 1.0; greedy decoding degenerates into verbatim loops (repetition-4 of 0.78–0.85), which is expected at this model size and token budget.
Trained on roughly 500M tokens each — small by current standards. These are base language models with no instruction tuning, alignment or safety filtering. They reproduce the distribution of their web-and-manual training corpora, including its biases. Context is capped at 512 tokens by the learned positional table.