Downloads · 30 days
31
21% of all-time downloads
kubataba/kbd-ru-opus
kbd-ru-opus is a translation model from kubataba. Use it when you need text moved from one language to another. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
Fine-tuned MarianMT model for Kabardian (East Circassian) to Russian translation.
Downloads · 30 days
31
21% of all-time downloads
All-time downloads
145
Public
Parameters
76.2M
612 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors305 MB · 99%
From the Hugging Face model README
Fine-tuned MarianMT model for Kabardian (East Circassian) to Russian translation.
This model translates from Kabardian to Russian. Kabardian is an endangered Northwest Caucasian language with approximately 500,000 speakers. It features complex polysynthetic morphology, 50+ consonants, and ergative-absolutive alignment.
Primary uses:
Limitations:
Dataset: adiga-ai/circassian-parallel-corpus
kbd_ru (Kabardian → Russian)base_model: Helsinki-NLP/opus-mt-en-ru
training_examples: 200,000
epochs: 7
batch_size: 32
learning_rate: 3e-4
optimizer: AdamW
max_sequence_length: 128
warmup_steps: 500
weight_decay: 0.01
framework: transformers 4.36.0
The model uses a special character mapping for training:
This ensures better tokenization compatibility with the MarianMT tokenizer.
Tested on 1,000 examples from adiga-ai/circassian-parallel-corpus:
| Metric | Score |
|---|---|
| BLEU | 28.13 |
| chrF | 50.07 |
| TER | 63.50 |
| Exact Match | 6.4% |
| Speed | 27.1 examples/sec |
| Avg Time | 37ms/example |
Test Configuration:
| Kabardian (Input) | Russian (Output) |
|---|---|
| Хабзэ зыхэмылъым жьантӀэр хуохъур кӀуапӀэ | У того, у кого нет обычаев, почетное место становится проходным местом. |
| КӀэщӀу жыпӀэмэ, я псэр зы чысэм илът. | Короче говоря, их душа лежала в одном кошельке. |
| Шухьэр Кесарие къалэм щынэсым, ӏэтащхьэм тхылъыр ир | Когда враги добрались до города Кесария, в главе... |
| щӀэныгъэм и унэтӀыныгъэщӀэм и ублакӀуэ | учитель нового направления науки |
| Сэтэней абы сагрисефэ щищӀауэ, нартыжьхэр щызэхуишэсауэ щызэхэст. | Сатаней сидела, собирав старинных нартов, сделав там сагрисефию. |
Note: The model successfully handles complex Kabardian morphology and preserves meaning in Russian translations.
pip install transformers torch sentencepiece
from transformers import MarianMTModel, MarianTokenizer
# Load model and tokenizer
model_name = "kubataba/kbd-ru-opus"
tokenizer = MarianTokenizer.from_pretrained(model_name)
model = MarianMTModel.from_pretrained(model_name)
# Translation function
def translate_kbd_to_ru(text):
# Preprocess: Ӏ → I for tokenization
processed_text = text.replace('Ӏ', 'I').replace('ӏ', 'I')
inputs = tokenizer(processed_text, return_tensors="pt", padding=True)
outputs = model.generate(**inputs, max_length=128, num_beams=4)
translation = tokenizer.decode(outputs[0], skip_special_tokens=True)
return translation
# Example
kabardian_text = "Сэлам!"
russian_text = translate_kbd_to_ru(kabardian_text)
print(f"KBD: {kabardian_text}")
print(f"RU: {russian_text}")
texts = [
"Уи пщэдджыжь фӀыуэ!",
"Сыт укъэпсэлъар?",
"Ди лъэпкъым и бзэр"
]
# Preprocess all texts
processed_texts = [t.replace('Ӏ', 'I').replace('ӏ', 'I') for t in texts]
inputs = tokenizer(processed_texts, return_tensors="pt", padding=True, truncation=True)
outputs = model.generate(**inputs, max_length=128, num_beams=4)
translations = [tokenizer.decode(out, skip_special_tokens=True) for out in outputs]
for src, tgt in zip(texts, translations):
print(f"KBD: {src} → RU: {tgt}")
This model contributes to digital language preservation for Kabardian, an endangered language.
Important considerations:
Kabardian (Adyghe-Kabardian, East Circassian) is a Northwest Caucasian language spoken by approximately 500,000 people in:
Linguistic features:
If you use this model in your research, please cite:
@misc{emkuzhev2025kbdru,
author = {Eduard Emkuzhev},
title = {Kabardian-Russian Neural Machine Translation Model},
year = {2025},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/kubataba/kbd-ru-opus}}
}
Please also cite the base model and dataset:
@misc{helsinki-nlp-opus-en-ru,
author = {Language Technology Research Group at the University of Helsinki},
title = {OPUS-MT English-Russian Translation Model},
year = {2020},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/Helsinki-NLP/opus-mt-en-ru}}
}
@dataset{qunash2025circassian,
author = {Anzor Qunash},
title = {Circassian-Russian Parallel Text Corpus v1.0},
year = {2025},
publisher = {adiga.ai},
url = {https://huggingface.co/datasets/adiga-ai/circassian-parallel-corpus}
}
For commercial licensing inquiries, please contact via email.
Model Card Authors: Eduard Emkuzhev
Last Updated: December 2025
Version: 1.0.1