Downloads · 30 days
186
100% of all-time downloads
nyobemedoc/spellmtei
spellmtei is a text generation model from nyobemedoc. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
The best corrector from SpellMtei, a spelling corrector for Manipuri (Meetei) in the Meetei Mayek script. google/byt5-small fine-tuned as a denoiser on synthetic supervision: clean Meetei Mayek text is corrupted by an…
Downloads · 30 days
186
100% of all-time downloads
All-time downloads
186
Public
Parameters
300M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 GB · 100%
From the Hugging Face model README
The best corrector from SpellMtei, a spelling corrector for Manipuri (Meetei)
in the Meetei Mayek script. google/byt5-small fine-tuned as a denoiser on
synthetic supervision: clean Meetei Mayek text is corrupted by an error model
grounded in the script's structure and in Meetei phonology, and the model learns
to reconstruct it.
Tokenizer-free (byte-level), so no vocabulary work for the script.
Anonymised for peer review. De-anonymised on acceptance.
| P | R | F0.5 | chrF | chrF++ | CER% | TER | FP-clean% | |
|---|---|---|---|---|---|---|---|---|
| Uncorrected input | – | – | – | 85.0 | 80.7 | 6.26 | 28.5 | 0.0 |
| M1 ByT5 (this model) | 35.4 | 41.3 | 36.5 | 91.4 | 88.4 | 2.86 | 16.4 | 60.0 |
Removes 54% of the character error and lifts chrF by 6.4 points. Like the other SpellMtei models it over-corrects already-correct input (it rewrites 60% of clean sentences) — raising the clean-sentence rate in training is the clearest next step.
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tok = AutoTokenizer.from_pretrained("nyobemedoc/spellmtei")
model = AutoModelForSeq2SeqLM.from_pretrained("nyobemedoc/spellmtei")
text = "ꯑꯩꯈꯣꯏꯒꯤ ꯅꯥꯛꯇꯥ ꯍꯧꯖꯤꯛ ꯂꯔꯦ" # dropped vowel sign
ids = tok(text, return_tensors="pt").input_ids
out = model.generate(ids, num_beams=4, max_length=1024)
print(tok.decode(out[0], skip_special_tokens=True)) # ... ꯍꯧꯖꯤꯛ ꯂꯩꯔꯦ
Feed one sentence at a time; NFC-normalise and collapse whitespace first.
google/byt5-small (~300M params), 6 epochs, one RTX 6000 Ada, bf16, Adafactor,
cosine schedule, max_len 1024, beam 4 at inference. Noise is resampled every
epoch; 20% of training pairs are left uncorrupted.
Error profile, frozen evaluation sets, pipeline code, and the other two
correctors (a from-scratch character edit tagger and a statistical noisy
channel): nyobemedoc/spellmtei (dataset).
To be finalised with the camera-ready release.