Downloads · 30 days
8
13% of all-time downloads
CidQuLimited/LazuriMT
LazuriMT is a translation model from CidQuLimited. Use it when you need text moved from one language to another. It is set up for peft. The card lists the license as gemma.
🇬🇧 English — LazuriMT is an open-source machine translation adapter for Turkish ↔ Laz (Lazuri), an endangered Kartvelian language spoken in northeastern Türkiye and parts of Georgia.
Downloads · 30 days
8
13% of all-time downloads
All-time downloads
63
Public
Repo size
913 MB
Likes
4
Public
Click a slice to open those files.
.safetensors587 MB · 95%
From the Hugging Face model README
🇬🇧 English — LazuriMT is an open-source machine translation adapter for Turkish ↔ Laz (Lazuri), an endangered Kartvelian language spoken in northeastern Türkiye and parts of Georgia.
🇹🇷 Türkçe — LazuriMT, Türkçe ↔ Lazca (Lazuri) arasında çeviri için açık kaynaklı bir adaptördür. Lazca, Türkiye'nin kuzeydoğusu ve Gürcistan'ın bazı bölgelerinde konuşulan, nesli tükenmekte olan bir Kartvel dilidir.
🌊 Lazuri — LazuriMT, Turkuli do Lazuri nenape şeni açikkaynaki tercüme modeli ren. Lazuri, Turkias do Gurcistanis na isinapunan Kartveluri nena ren.
LoRA adapter for Gemma 4 E4B. v0.2 research preview.
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "unsloth/gemma-4-e4b-it-unsloth-bnb-4bit"
model = AutoModelForCausalLM.from_pretrained(base, device_map="auto", load_in_4bit=True)
model = PeftModel.from_pretrained(model, "CidQuLimited/LazuriMT")
tok = AutoTokenizer.from_pretrained("CidQuLimited/LazuriMT")
def translate(text, to="lzz"):
prompt = (f"Translate this Turkish sentence into Laz (Lazuri):\n\n{text}"
if to == "lzz"
else f"Translate this Laz (Lazuri) sentence into Turkish:\n\n{text}")
inputs = tok.apply_chat_template(
[{"role": "user", "content": prompt}],
tokenize=True, add_generation_prompt=True, return_tensors="pt",
).to(model.device)
out = model.generate(
input_ids=inputs, max_new_tokens=128, do_sample=False,
no_repeat_ngram_size=3, repetition_penalty=1.15, num_beams=4,
)
return tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True).strip()
print(translate("Su içmek istiyorum."))
Pin to a specific release with revision="v0.2" (or "v0.1" for the older one):
model = PeftModel.from_pretrained(model, "CidQuLimited/LazuriMT", revision="v0.2")
chrF computed on 200 held-out TR→LZ pairs from the corpus's test split (5%), with beam-search decoding (no_repeat_ngram_size=3, repetition_penalty=1.15, num_beams=4).
| Version | chrF (TR→LZ) | Notes |
|---|---|---|
| baseline Gemma 4 E4B (no adapter) | ≈ 0 | does not translate Laz |
| v0.1 | 24.66 | LoRA r=32, 10,500 masked-loss steps (~2.15 epochs), Kaggle T4 |
| v0.2 (this release) | 26.97 | LoRA r=64, 18,000 steps (3 epochs), A100, cosine-restart LR, 3× dialect upweight |
For context, chrF roughly maps:
LazuriMT v0.2 is in the "readable but flawed" range — a real but early baseline for a language with almost no prior MT.
unsloth/gemma-4-e4b-it-unsloth-bnb-4bit (Gemma 4 E4B, pre-quantized to 4-bit)r=64, α=64, dropout 0lr=2e-4, cosine-with-restarts (2 cycles), warmup_ratio 0.03, bf16[Laz dialect: X] label in the prompt — but it did not meaningfully change behavior. The likely cause: even at 3×, dialect-tagged pairs are only ~9 % of the training mix, so the model defaults to general-form Laz. v0.3 will try a dialect-balanced sampler (equal exposure per dialect rather than blunt upweighting) plus additional dialect-tagged parallel data.The adapter is derivative work of Gemma 4 and inherits the Gemma Terms of Use — commercial-friendly but with acceptable-use restrictions. Downstream users must comply with Gemma's terms.
The training corpus mixes open-license sources (Wikipedia CC-BY-SA, Mozilla Common Voice CC-BY, GPL-3.0 Wiktionary, public-domain Lazuri Paramitepe 1982) with academically-attributed sources used under fair-use for endangered-language research. The adapter weights are released for research and community use under these combined terms.
@misc{lazurimt2026,
title = {LazuriMT: A Turkish-Laz Machine Translation Adapter for an Endangered Kartvelian Language},
author = {Yavuz Selimhan Kaya},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/CidQuLimited/LazuriMT}},
note = {v0.2 research preview, chrF 26.97 on 200 TR→LZ test pairs}
}
The Lazuri community: İsmail Avcı Bucaklişi, Hasan Uzunhasanoğlu (Lazuri.Com), Ali İhsan Aksamaz, Özlem Durmaz (translator of Anadolu Dillerinde Küçük Prens), the Laz Institute, contributors to lazcasozluk.org and the Lazuri Wiktionary GitHub project, the Ministry of National Education of Türkiye, the broader Laz language preservation community, and every Laz speaker who has kept this language alive.
Tools: Unsloth for the QLoRA training stack, Google's Gemma 4 as the base model, the Hugging Face ecosystem, and Kaggle for the GPU compute.
Full reproduction code, training data sources, and iteration history: https://github.com/CidQu/lazca_ai