Downloads · 30 days
0
Prieyan/EnglishToHindi
EnglishToHindi is a machine learning model from Prieyan. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A small encoder-decoder transformer that translates English into Hindi, trained from scratch on a ~2,000-sentence parallel corpus.
Downloads · 30 days
0
Access
Public
Updated Aug 17, 2026
Repo size
40.7 MB
Likes
0
Public
Click a slice to open those files.
.csv40.7 MB · 100%
From the Hugging Face model README
A small encoder-decoder transformer that translates English into Hindi, trained from scratch on a ~2,000-sentence parallel corpus.
Vanilla "Attention Is All You Need" architecture — no pretrained weights, no frameworks beyond PyTorch, everything from the tokenizer to the BLEU score written out in full.
| Parameters | 6.55M |
| Encoder | 3 layers |
| Decoder | 3 layers (self-attention + cross-attention) |
| Hidden dim | 256 |
| Heads | 4 |
| FFN | 1024 (ReLU) |
| Vocab | 4,000 BPE, shared across both languages |
| Max sequence | 64 tokens |
| Norm | LayerNorm (pre-norm) |
| Position | Fixed sinusoidal |
| Embeddings | Tied across encoder, decoder and output head |
| Decoding | Greedy and beam search |
The encoder reads the English sentence; the decoder generates Hindi, attending to its own output so far and, through cross-attention, to the whole source.
pip install -r requirements.txt
python main.py prepare
Splits data/translation.csv into train/val/test and trains the shared BPE
tokenizer. The split is grouped by English sentence, so no source sentence
appears in two splits, and the tokenizer only ever sees the training half.
python main.py train
~60 seconds on a GPU, a few minutes on CPU. The best checkpoint by validation
loss lands in checkpoints/best_model.pt.
python main.py translate # interactive
python main.py translate -s "Are you at home?" # one sentence
python main.py translate -s "Good morning" --beam 1 # greedy
python main.py benchmark # writes benchmarks/benchmark_<timestamp>.txt
python main.py benchmark --cpu # force CPU
python main.py benchmark --cpu --threads 4
python main.py benchmark --beam 1 4 8 # compare decoding strategies
python main.py benchmark --json out.json # raw metrics too
Held-out test split (90 source sentences), scored against every reference translation available for each sentence:
| Greedy | Beam (4) | |
|---|---|---|
| BLEU | 5.38 | 6.90 |
| chrF | 24.86 | 25.55 |
| Exact match | 1.1% | 1.1% |
| Speed | 59.3 sent/s | 25.7 sent/s |
Teacher-forced on the same split: loss 2.55, perplexity 12.84, top-1 token accuracy 51.2%, top-5 74.7%.
These numbers are low, and honestly so. 1,797 training pairs is a tiny corpus for translation — the model reliably picks up sentence shape, question particles and common phrases, then invents vocabulary it never saw enough of:
EN Are you at home?
OUT तुम घर हो क्या?
REF तुम घर पे हो क्या?
EN Bad news travels quickly.
OUT बच्चे सेवार को साथ-कार करो।
REF बुरी खबर तेज़ी से फैलती है।
Validation loss bottoms out around step 1,200 and rises afterwards while training loss keeps falling — classic overfitting on a small corpus. More parallel data is the only real fix; label smoothing, dropout 0.2 and tied embeddings already do what regularisation can.
Both scores are implemented in slm/metrics.py, not imported:
Where an English sentence has several Hindi translations in the corpus, all of them count as references.
data/translation.csv English–Hindi parallel corpus
main.py CLI entry point
slm/config.py Model and training hyperparameters
slm/data.py Corpus loading, cleaning, splitting
slm/tokenizer.py Shared byte-level BPE tokenizer
slm/prepare_data.py Split + tokenizer pipeline
slm/dataset.py Tokenized pairs, padding, batching
slm/model.py Attention, encoder, decoder, decoding
slm/train.py Training loop
slm/translate.py Inference API and CLI
slm/metrics.py BLEU and chrF
slm/benchmark.py Benchmark suite and report writer