Downloads · 30 days
18
47% of all-time downloads
vivekharry/bohok-indic-en
bohok-indic-en is a translation model from vivekharry. Use it when you need text moved from one language to another. It is set up for transformers. The card lists the license as mit.
Marian MT (Helsinki-NLP/opus-mt-mul-en) fine-tuned one epoch on CPU so you can load a real checkpoint with transformers and translate Kokborok, Bengali, and Marathi into English.
Downloads · 30 days
18
47% of all-time downloads
All-time downloads
38
Public
Parameters
77.1M
310 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors308 MB · 99%
From the Hugging Face model README
vivekharry/bohok-indic-en)Marian MT (Helsinki-NLP/opus-mt-mul-en) fine-tuned one epoch on CPU so you can load a real checkpoint with transformers and translate Kokborok, Bengali, and Marathi into English.
This is not a universal translator and not Google Translate. It is a small many-to-English model adapted on:
| Language | Training pairs | Source |
|---|---|---|
Kokborok (trp) | ~10k | SMOL sentences + GATITOS + SMOL-doc (sdmy / Google SMOL) |
Bengali (bn) | 12k | OPUS-100 bn-en |
Marathi (mr) | 12k | OPUS-100 en-mr (Marathi side → English) |
Held-out split: 4% (~1360 sentences). Train set: 32,657. One CPU epoch (4,083 steps, 42m). Train loss 2.49, eval loss 2.05.
Held-out examples after training (not cherry-picked UI strings):
| Lang | Source (truncated) | Model |
|---|---|---|
| bn | আমার দিকে তাকাও. | Look at me. |
| bn | বাসায় এটি পালন করবেন না! | Don't do this in the house! |
| mr | टॉम जेवतोय. | Tom is eating. |
| mr | 90 डिग्री घड्याळीचे उलट दिशेने फिरविले | Rotated 90 degrees counter-clockwise |
| trp | (SMOL sentences) | Often topical English, weaker than bn/mr — 10k Kokborok pairs vs 24k OPUS |
Kokborok is the low-resource track. Bengali/Marathi inherit OPUS-100. This is not Google Translate.
The tokenizer gained three special tokens. Always prefix the source:
>>trp<< <kokborok text>
>>bn<< <bengali text>
>>mr<< <marathi text>
Kokborok training data is mostly Latin / Roman SMOL spelling, plus some Bangla-script items from GATITOS. Roman Kokborok works better than free dialect you invent.
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
repo = "vivekharry/bohok-indic-en"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSeq2SeqLM.from_pretrained(repo)
def to_en(text, lang="trp"):
ids = tok(f">>{lang}<< {text}", return_tensors="pt", truncation=True, max_length=128)
out = model.generate(**ids, max_new_tokens=96, num_beams=4)
return tok.decode(out[0], skip_special_tokens=True)
print(to_en("আমার দিকে তাকাও.", "bn"))
print(to_en("टॉम जेवतोय.", "mr"))
CLI from this repo:
python scripts/translate.py --model vivekharry/bohok-indic-en --lang bn "আজ আকাশটা খুব সুন্দর।"
python scripts/prepare_bohok_data.py
python scripts/train_bohok_mt.py --epochs 1 --bs 8 --out checkpoints/bohok-indic-en
python scripts/eval_and_push.py --repo vivekharry/bohok-indic-en
Author: Vivek Das (vivekharry).