Downloads · 30 days
17
5% of all-time downloads
leks-forever/mt5-base
mt5-base is a translation model from leks-forever. Use it when you need text moved from one language to another. It is set up for transformers. The card lists the license as apache-2.0.
This version of the Google T5-Base model has been fine-tuned on a bilingual dataset of Russian and Lezgian sentences to improve translation quality in both directions (from Russian to Lezgian and from Lezgian to Russi…
Downloads · 30 days
17
5% of all-time downloads
All-time downloads
311
Public
Parameters
582M
2.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 99%
From the Hugging Face model README
This version of the Google T5-Base model has been fine-tuned on a bilingual dataset of Russian and Lezgian sentences to improve translation quality in both directions (from Russian to Lezgian and from Lezgian to Russian). The model is designed to provide accurate and high-quality translations between these two languages.
"translate Russian to Lezghian: " - Ru-Lez
"translate Lezghian to Russian: " - Lez-Ru
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
model = AutoModelForSeq2SeqLM.from_pretrained("leks-forever/mt5-base")
tokenizer = AutoTokenizer.from_pretrained("leks-forever/mt5-base")
def predict(text, prefix, a=32, b=3, max_input_length=1024, num_beams=1, **kwargs):
inputs = tokenizer(prefix + text, return_tensors='pt', padding=True, truncation=True, max_length=max_input_length)
result = model.generate(
**inputs.to(model.device),
max_new_tokens=int(a + b * inputs.input_ids.shape[1]),
num_beams=num_beams,
**kwargs
)
return tokenizer.batch_decode(result, skip_special_tokens=True)
sentence: str = "Римдин аскерар ва гьакӀни чӀехи хахамрини фарисейри ракъурнавай нуькерар Ягьуд галаз багъдиз атана. Абурув виридав яракьар, чирагъар ва шемгьалар гвай."
translation = predict(sentence, prefix="translate Lezghian to Russian: ")
print(translation)
# ['Когда римские воины и вожди, а также главные священнослужители и блюстители Закона пришли в Иудею, они дали ему вооружённые оружие, браслеты и серьги.']
The model was fine-tuned on the bible-lezghian-russian dataset, which contains 13,800 parallel sentences in Russian and Lezgian. The dataset was split into three parts: 90% for training, 5% for validation, and 5% for testing.
The evaluation was conducted on the val set of the bible-lezghian-russian dataset, consisting of 5% of the total 13,800 parallel sentences.
The evaluation considered translations in both directions:
The following metrics were used to evaluate the model’s performance:
These results indicate that the model can produce accurate translations for both language pairs. However, there are plans to improve the model further by conducting parallel alignment of the corpora to refine the sentence pair matching. Additionally, efforts will be made to collect more training data to enhance the model's performance, especially in handling more diverse and complex linguistic structures.
<!-- ## More Information [optional] [More Information Needed] ## Model Card Authors [optional] [More Information Needed] ## Model Card Contact [More Information Needed] -->