Downloads · 30 days
29
6% of all-time downloads
islomov/rubai-corrector-base
rubai-corrector-base is a text generation model from islomov. Use it when you need the model to write or continue text. It is set up for transformers.
Base ByT5 correction checkpoint for building task-specific Rubai correctors.
Downloads · 30 days
29
6% of all-time downloads
All-time downloads
494
Public
Parameters
300M
1.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors1.2 GB · 100%
From the Hugging Face model README
Base ByT5 correction checkpoint for building task-specific Rubai correctors.
This is the foundation model of the rubai-corrector line. It is meant to be fine-tuned for a concrete demand:
If you want a ready-to-use ASR display model, use rubai-corrector-transcript-uz. If you want the OCR-specialized old-books model, use rubai-corrector-ocr-books-uz. This package is the base for further adaptation.
| Model | Use Case |
|---|---|
| rubai-corrector-base (this model) | Fine-tuning base for new correction tasks |
| rubai-corrector-transcript-uz | Ready-to-use transcript display normalization |
| rubai-corrector-ocr-books-uz | OCR correction for old Uzbek books |
Both models share the same ByT5 architecture. The transcript model is fine-tuned from this base for ASR display text.
The model uses the correct: instruction prefix.
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
model_path = "islomov/rubai-corrector-base"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForSeq2SeqLM.from_pretrained(model_path)
text = "men ozim kordim"
inputs = tokenizer([f"correct: {text}"], return_tensors="pt", padding=True)
output_ids = model.generate(**inputs, max_new_tokens=128)
prediction = tokenizer.batch_decode(output_ids, skip_special_tokens=True)[0]
print(prediction)
Expected output:
Men o'zim ko'rdim
For a local runnable example suite, see test_model.py.
These are real outputs from this packaged checkpoint.
Input: telefon rqami qaysi
Output: Telefon raqami qaysi
Input: men ozim kordim
Output: Men o'zim ko'rdim
Input: togri yoldan boring
Output: To'g'ri yo'ldan boring
Input: rnen universitetda oqiyrnan
Output: Men universitetda o'qiyman
Input: bu juda rnuhirn masala
Output: Bu juda muhim masala
Input: narxi yigirma besh ming so'm
Output: Narxi 25 000 so'm
Input: uchrashuv o'n beshinchi yanvar kuni
Output: Uchrashuv 15-yanvar kuni
Input: men segodnya bozorga bordim
Output: Men сегодня bozorga bordim
Input: privet kak делa
Output: Привет как дела
This package includes a standalone fine-tuning script:
It keeps the same core training behavior as the original project line:
correct: T5ForConditionalGenerationinput -> output pairsExample:
python finetune.py \
--model-path rubai/rubai-corrector-base \
--train-file ./data/train.jsonl \
--eval-file ./data/valid.jsonl \
--output-dir ./outputs/my-domain-corrector \
--learning-rate 5e-5 \
--num-train-epochs 2 \
--per-device-train-batch-size 16 \
--gradient-accumulation-steps 4 \
--max-source-length 512 \
--max-target-length 512 \
--bf16
Training data is JSONL. Each line must contain:
input: noisy or source textoutput: target corrected textExample:
{"input":"men ozim kordim","output":"Men o'zim ko'rdim"}
{"input":"narxi yigirma besh ming so'm","output":"Narxi 25 000 so'm"}
{"input":"rnen universitetda oqiyrnan","output":"Men universitetda o'qiyman"}
{"input":"men segodnya bozorga bordim","output":"Men сегодня bozorga bordim"}
A tiny sample file is included here:
You can point finetune.py either to a JSONL file directly or to a directory containing data.jsonl.
This model starts from google/byt5-small and was built with a 3-stage curriculum on Uzbek text correction data.
The foundation stage used ~1,000,000 synthetic correction pairs generated from Uzbek text with transformations such as:
h/x swapsStage 2 added ~408,000 curated rows covering:
h/x restorationStage 3 used ~32,000 rows for fine-grained behavior tuning:
T5ForConditionalGeneration with ByT5 tokenizerSpecial thanks to Davron Ibrokhimov for sponsoring this work and making it possible to keep these models open.
Thank you to the community that supports Uzbek language technology. In particular:
Thanks to Arofat, Gulimshaxnoz, and many others who contributed in ways big and small. The list is too long to fit here, but every contribution matters and is appreciated.
Support my works and open-source movement: https://tirikchilik.uz/islomovs