Downloads · 30 days
0
meet5568/lma_models
lma_models is a text generation model from meet5568. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This repository holds both phases of the project: the two sentencepiece tokenizers (Phase 1) and the two pretrained decoder-only Transformers that use them (Phase 2).
Downloads · 30 days
0
Access
Public
Updated Sep 15, 2026
Repo size
14 GB
Likes
0
Public
Click a slice to open those files.
.pt5.1 GB · 100%
From the Hugging Face model README
This repository holds both phases of the project: the two sentencepiece tokenizers (Phase 1) and the two pretrained decoder-only Transformers that use them (Phase 2).
Two independent sentencepiece tokenizers, one per language, trained from scratch for a pair of ~25M-parameter decoder-only Transformers.
They share nothing — not merges, not pieces, not a vocabulary file. Hindi and Nepali both use the Devanagari block (U+0900–U+097F), so keeping the two corpora and the two vocabularies separate is the central constraint of the project rather than an afterthought.
| language | algorithm | vocab | trained on | corpus tokens |
|---|---|---|---|---|
| Hindi | unigram | 10,000 | 2,016,377 of 4,032,755 lines (50%, sampled at random) | 655.2M |
| Nepali | unigram | 10,000 | 2,539,821 of 5,079,643 lines (50%, sampled at random) | 558.9M |
Twenty models were compared: five vocabulary sizes (8k, 10k, 12k, 14k, 16k) across two algorithms (BPE, unigram), for both languages. Every model read the same 10% random sample of its language's training split, so differences between them are differences between models rather than between samples.
Unigram beat BPE at every vocabulary size in both languages — ten paired comparisons, no exceptions, by 1.2–2.7% fertility.
Vocabulary 10,000 was selected over 16,000 despite 16k tokenizing better.
The embedding matrix is vocab_size × 512 parameters against a 25M budget, so
16k spends 33% of the whole model on a lookup table while 10k spends 20%. The
selection rule is the smallest vocabulary whose fertility is within 8% of the
best, which trades ~5–7% fertility for ~3.1M parameters returned to the
transformer layers.
Hindi — fertility 1.3250 tokens/word, 3.7749 chars/token, 0 UNK, 0.289% byte-fallback, 201 unused pieces. Nepali — fertility 1.4620 tokens/word, 4.5208 chars/token, 0 UNK, 0.178% byte-fallback, 239 unused pieces.
UNK is impossible. byte_fallback=True decomposes any unseen character
into byte tokens, so the UNK count is zero by construction rather than by luck.
| Hindi | Nepali | |
|---|---|---|
| final corpus | 5,094,185 docs / 2.588B chars | 6,755,888 docs / 2.513B chars |
| training split | 3,564,832 docs / 1.798B chars | 4,726,832 docs / 1.752B chars |
Splits are document-level, stratified by source, 70/15/15 by characters, with one whole source held out per language as an unseen-domain test set. The tokenizers saw the training split only.
The corpora are at meet5568/lma_datasets.
import sentencepiece as spm
from huggingface_hub import hf_hub_download
path = hf_hub_download("meet5568/lma_models", "tokenizer/hindi/hi_tokenizer.model")
sp = spm.SentencePieceProcessor(model_file=path)
print(sp.encode("भारत एक विशाल देश है।", out_type=str))
Unigram's memory scales with total corpus length — it builds a suffix array over every character — at roughly 10.6 GB of RAM per GB of text, measured. The full training split would need about 49 GB, so the final unigram models were trained on 50% of it. A control experiment found fertility differing in the fourth decimal place between 81.6% and 100% of the lines, so this is a hardware limit rather than a quality one, but it is stated rather than implied.
Two decoder-only Transformers, one per language, pretrained from scratch on the corpora above. They share an implementation and nothing else: separate data, separate tokenizers, separate vocabularies, separate weights.
Written from PyTorch primitives only — nn.Linear, nn.Embedding,
nn.LayerNorm, nn.Dropout. No nn.Transformer*, no nn.MultiheadAttention,
no F.scaled_dot_product_attention; the attention is written out.
| Model H (Hindi) | Model L (Nepali) | |
|---|---|---|
| parameters | 26,423,040 | 26,423,040 |
| layers x d_model | 4 x 640 | 4 x 640 |
| heads | 8 x 80 | 8 x 80 |
| d_ff | 2560 (GELU) | 2560 (GELU) |
| context | 512 | 512 |
| positions | learned absolute | learned absolute |
| embeddings | tied input/output | tied input/output |
| training tokens | 458,620,928 (1 epoch) | 394,461,184 (1 epoch) |
| steps | 6,998 | 6,019 |
| split | Hindi ppl | Hindi bpb | Nepali ppl | Nepali bpb |
|---|---|---|---|---|
| validation | 25.63 | 0.4644 | 36.55 | 0.4310 |
| test | 25.66 | 0.4637 | 37.38 | 0.4374 |
| unseen domain | 25.60 | 0.4756 | 36.15 | 0.4206 |
Compare the two models on bits per byte, not perplexity. Perplexity is measured per token and the two tokenizers differ (10.10 bytes per token for Hindi against 11.94 for Nepali), so Nepali is charged for packing more text into each prediction. Bits per byte divides by the raw UTF-8 length and reverses the ranking. The same caution applies to comparing these numbers with a model trained on a different corpus with a different tokenizer.
A full factorial sweep per language — learning rate x dropout x shape x context = 24 trials each, 48 in total, every trial 400 steps at 26.2M tokens. All 48 completed and none diverged. Both languages independently chose the same configuration. Learning rate mattered most (1.2e-3 beat 6e-4 everywhere); width beat depth at a fixed parameter budget; context 512 beat 1024, because documents average 82-128 tokens; dropout 0.0 beat 0.1, as a single epoch has nothing to overfit.
AdamW (betas 0.9/0.95, weight decay 0.1 on matrices only), 400-step warmup then cosine from 0.0012 to 1.2e-04, 65,536 tokens per optimisation step, gradient clipping 1.0, AMP fp16. One epoch, so every token is seen exactly once.
import torch
from huggingface_hub import hf_hub_download
path = hf_hub_download("meet5568/lma_models", "pretrained/hindi/best.pt")
# weights_only=False is required: the checkpoint carries optimizer, scaler and
# RNG state so a run can resume exactly, and torch 2.6+ refuses those under the
# default weights_only=True.
ckpt = torch.load(path, map_location="cpu", weights_only=False)
print(ckpt["model_config"]) # architecture
print(ckpt["step"], ckpt["best_val_loss"]) # where it stopped
state = ckpt["model"] # the weights themselves
Each checkpoint holds the model weights, optimizer state, gradient-scaler state, the training step, the epoch and window cursor, and both configs — so a run resumes exactly where it stopped rather than approximately.
run.json beside each checkpoint repeats the architecture and the training
hyperparameters in plain text and adds the full loss curve, so the settings can
be read without loading 317 MB of weights.
The model code is not on the hub. ckpt["model"] is a plain state_dict for
the implementation in
the project repository, whose
common/model_train/model/ builds the architecture from model_config.
Both models are under-trained. At ~26.4M parameters the Chinchilla-optimal budget is roughly 528M tokens; one epoch gives 17.4 tokens per parameter for Hindi and 14.9 for Nepali, against the ~20 rule of thumb. Both losses were still falling when the epoch ended.
The sweep ranked 400-step behaviour, and its winner sits on the boundary of the grid on three of four axes, so those values are the best tested rather than the best. One seed per configuration.
The unseen-domain split did not work as a robustness test — both models score better on it than in domain by perplexity, which says that domain is simpler than the corpus average rather than that the models generalise well.
The Phase 2 checkpoints above, finetuned on a synthetic comparative reasoning corpus generated separately for each language, at five sample sizes (5k, 10k, 15k, 30k, 50k, 100k). Every size is an independent run from the pretrained checkpoint: 5 epochs, batch 32, peak learning rate 1e-4 with warmup and cosine decay, loss on the answer tokens only, best validation checkpoint kept. Same architecture, tokenizer and 26.4M parameters as Phase 2.
Four entity domains (people, cities, items, vehicles), ten attributes (age,
height, score, weight, income, experience, temperature, population, price,
speed), four families (compare, transitive, superlative, ordering), six phrasings
each — every word counted in the pretraining corpus first.
100,000 training examples; smaller sizes are prefixes. Corpora:
meet5568/lma_datasets under
reasoning/.
| Split | Entity names | Numeric values | Prompt templates | Chain length |
|---|---|---|---|---|
| Test A | Unseen | Unseen | Seen | 2 |
| Test B (hardest) | Unseen | Unseen | Unseen | 2 |
| Test C | Unseen | Seen range | Unseen | 2 |
| Test D (control) | Seen | Seen range | Seen | 2 |
| Test E (longer chains) | Seen | — | Seen | 3–4 |
600 items per split; 0 duplicate prompts across all splits.
Each cell is Choice / Gen (Invalid): choice accuracy ranks the legal answers by log-probability; generation accuracy is greedy decoding, first word (punctuation stripped) exact; invalid is the share of generations that are not a legal answer. PPL and BPB are measured on the held-out pretraining corpus.
| Sample size | Steps | Train time | Best val loss | Test A | Test B | Test C | Test D | Test E | PPL | BPB |
|---|---|---|---|---|---|---|---|---|---|---|
| Pretrained | 0 | — | — | 32.7% / 0.0% (100.0%) | 33.7% / 0.0% (99.7%) | 36.2% / 0.0% (100.0%) | 30.5% / 0.0% (99.3%) | 51.0% / 0.0% (100.0%) | 25.66 | 0.4637 |
| 5k (5,000) | 780 | 29s | 1.3992 | 37.8% / 22.2% (47.5%) | 39.7% / 21.8% (46.7%) | 40.5% / 24.0% (45.5%) | 40.5% / 40.7% (0.3%) | 47.0% / 47.0% (0.0%) | 36.42 | 0.5138 |
| 10k (10,000) | 1,560 | 60s | 1.2472 | 36.8% / 20.3% (50.0%) | 36.5% / 19.3% (50.0%) | 40.0% / 21.3% (48.8%) | 39.3% / 39.3% (0.2%) | 45.3% / 45.3% (0.0%) | 52.73 | 0.5667 |
| 15k (15,000) | 2,340 | 92s | 1.1723 | 40.0% / 23.5% (48.5%) | 40.8% / 24.2% (49.0%) | 39.3% / 22.7% (47.0%) | 39.5% / 40.7% (0.5%) | 54.0% / 54.0% (0.0%) | 70.59 | 0.6083 |
| 30k (30,000) | 4,685 | 183s | 1.0630 | 37.2% / 21.5% (50.0%) | 39.3% / 23.3% (50.0%) | 39.3% / 22.8% (48.8%) | 37.8% / 38.3% (0.0%) | 57.5% / 57.5% (0.0%) | 188.73 | 0.7489 |
| 50k (50,000) | 7,810 | 304s | 1.0073 | 40.2% / 23.5% (50.0%) | 39.2% / 23.8% (50.0%) | 41.3% / 23.3% (48.8%) | 40.0% / 40.3% (0.3%) | 56.7% / 56.7% (0.0%) | 498.58 | 0.8877 |
| 100k (100,000) | 15,625 | 10.2 min | 0.7857 | 47.0% / 30.5% (47.0%) | 47.7% / 33.7% (47.3%) | 52.2% / 36.7% (45.2%) | 64.8% / 63.8% (1.8%) | 79.7% / 79.7% (0.0%) | 1720.75 | 1.0647 |
| chance (choice) | 35.6% | 35.4% | 35.4% | 35.6% | 45.0% | |||||
| majority answer | 44.5% | 44.5% | 44.3% | 47.0% | 54.8% |
| Sample size | Steps | Train time | Best val loss | Test A | Test B | Test C | Test D | Test E | PPL | BPB |
|---|---|---|---|---|---|---|---|---|---|---|
| Pretrained | 0 | — | — | 35.8% / 0.0% (100.0%) | 38.5% / 0.0% (100.0%) | 36.3% / 0.0% (100.0%) | 33.2% / 0.0% (100.0%) | 51.5% / 0.0% (100.0%) | 37.38 | 0.4374 |
| 5k (5,000) | 780 | 25s | 1.2820 | 39.5% / 23.0% (50.0%) | 39.3% / 22.0% (49.8%) | 36.0% / 20.3% (48.7%) | 41.7% / 33.8% (20.8%) | 50.2% / 48.7% (0.0%) | 58.26 | 0.4910 |
| 10k (10,000) | 1,560 | 50s | 1.1704 | 37.5% / 20.5% (50.0%) | 39.8% / 21.2% (50.0%) | 37.5% / 21.7% (48.8%) | 38.2% / 33.5% (17.8%) | 51.2% / 50.8% (0.0%) | 94.50 | 0.5495 |
| 15k (15,000) | 2,340 | 75s | 1.0940 | 36.3% / 21.0% (50.0%) | 35.8% / 18.0% (50.0%) | 35.0% / 20.0% (48.8%) | 40.2% / 38.0% (9.2%) | 50.2% / 50.2% (0.0%) | 150.04 | 0.6053 |
| 30k (30,000) | 4,685 | 151s | 0.9765 | 36.5% / 20.8% (50.0%) | 35.3% / 18.2% (50.0%) | 34.7% / 20.0% (48.8%) | 43.0% / 41.3% (3.8%) | 49.7% / 49.0% (0.0%) | 788.17 | 0.8057 |
| 50k (50,000) | 7,810 | 251s | 0.8978 | 38.0% / 22.0% (50.0%) | 36.8% / 20.7% (50.0%) | 34.5% / 20.3% (48.8%) | 42.5% / 40.0% (7.3%) | 49.3% / 49.5% (0.0%) | 9686.89 | 1.1087 |
| 100k (100,000) | 15,625 | 502s | 0.6837 | 47.0% / 31.5% (50.0%) | 45.5% / 30.7% (50.0%) | 49.2% / 34.7% (48.8%) | 58.0% / 55.0% (7.2%) | 50.3% / 49.7% (0.0%) | 23869.70 | 1.2177 |
| chance (choice) | 37.5% | 37.4% | 37.4% | 37.5% | 50.0% | |||||
| majority answer | 44.5% | 44.5% | 44.3% | 44.3% | 51.5% |
Hindi. Pretrained, the model is close to chance (36.8% choice against 37.4%) and produces a legal answer in generation 0.2% of the time.
transitive items (-12.3 points), with 54.8% for always giving E's commonest answer.superlative items — whose answer is an entity name — are invalid 97.3% of the time on split A (held-out names) against 3.3% on D (training names): the model writes a name it saw in training rather than the new name in the prompt.Nepali. Pretrained, the model is close to chance (39.1% choice against 40.0%) and produces a legal answer in generation 0.0% of the time.
transitive items stay at 46.7% (chance 50.0%), so this model never learned to chain two facts and E (50.3%) cannot show a chain-length effect; its gains come from the other question types.superlative items — whose answer is an entity name — are invalid 100.0% of the time on split A (held-out names) against 14.7% on D (training names): the model writes a name it saw in training rather than the new name in the prompt.import torch
from huggingface_hub import hf_hub_download
path = hf_hub_download("meet5568/lma_models", "finetuned/hindi/100k/best.pt")
ckpt = torch.load(path, map_location="cpu", weights_only=False)
Prompt with the question followed by the answer marker
(जवाब for Hindi, जवाफ
for Nepali) and read the next word.
The task is synthetic and templated. Finetuning on it degrades general language modelling, visible in the PPL column, and the models almost never generate a held-out entity name.