Downloads · 30 days
9
35% of all-time downloads
worldboss/en-ee-from-scratch
en-ee-from-scratch is a translation model from worldboss. Use it when you need text moved from one language to another. It is set up for pytorch. The card lists the license as cc-by-nc-4.0.
From-scratch English → Ewe encoder–decoder (Pre-LN, 6×512, vocab 32000). This is not an NLLB / Transformers AutoModelForSeq2SeqLM checkpoint. Load best.pt with translate.py from the training repo.
Downloads · 30 days
9
35% of all-time downloads
All-time downloads
26
Public
Repo size
242 MB
Likes
0
Public
Click a slice to open those files.
.pt242 MB · 99%
From the Hugging Face model README
From-scratch English → Ewe encoder–decoder (Pre-LN, 6×512, vocab 32000).
This is not an NLLB / Transformers AutoModelForSeq2SeqLM checkpoint. Load best.pt
with translate.py from the training repo.
Training data includes Ghana farmer Q&A, NLLB eng_Latn-ewe_Latn, and verse-aligned
English–Ewe Bible from ghananlpcommunity/ghana-corpus.
Farmer Q&A, NLLB, and the Bible corpus are CC-BY-NC.
huggingface-cli download worldboss/en-ee-from-scratch --local-dir en-ee-from-scratch
python translate.py --checkpoint en-ee-from-scratch/best.pt --text "How do I plant maize?"
from pathlib import Path
import torch
from translate import load_checkpoint, translate
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model, tokenizer, cfg = load_checkpoint(Path("en-ee-from-scratch/best.pt"), device)
print(translate(model, tokenizer, cfg, "How do I plant maize?", device))
The Hub snapshot ships best.pt plus the BPE tokenizer (tokenizer.json).
translate.py resolves tokenizer_dir="." next to the checkpoint.
Sample prompts:
Compare with the NLLB fine-tune using python compare_models.py in the training repo.