Downloads · 30 days
0
Anadilorg/LazuriMT
LazuriMT is a machine learning model from Anadilorg. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
\\\\\\\\\\\\\\Bu model şu anda yayına kapalıdır.\\\\\\\\\\\\\\ \\\\\\\\\\\\\\This model is currently not publicly available.\\\\\\\\\\\\\\ Güncellemeler için \\\\\\\\\\\\\\\\\\\\\anadil.org\\\\\\\\\\\\\\ adresini taki…
Downloads · 30 days
0
Access
Public
Updated Aug 25, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.png76.2 KB · 88%
From the Hugging Face model README
\\\*\\\*Bu model şu anda yayına kapalıdır.\\\*\\\*
\\\*\\\*This model is currently not publicly available.\\\*\\\*
Güncellemeler için \\\*\\\*\\\[anadil.org](https://anadil.org)\\\\\\\*\\\\\\\* adresini takip edin.
For updates, follow \\\*\\\*\\\[anadil.org](https://anadil.org)\\\\\\\*\\\\\\\*.
---
LazuriMT, Google'ın TranslateGemma-4B-IT modelini temel alarak geliştirilen, Lazca (ISO 639-3: lzz) desteği eklenmiş deneysel bir çokdilli çeviri modelidir. Türkçe, İngilizce, İspanyolca, Almanca, Rusça ve Arapça'dan Lazca'ya yüksek kaliteli çeviriler üretir.
LazuriMT is an experimental multilingual translation model based on Google's TranslateGemma-4B-IT, with Lazca (ISO 639-3: lzz) support added via LoRA fine-tuning. It produces high-quality translations from Turkish, English, Spanish, German, Russian, and Arabic into Lazca.
---

| Base Model | google/translategemma-4b-it (Gemma 3 family, 4B parameters, 55 languages) |
| Adapter Type | LoRA (r=32, alpha=64, dropout 0.05) |
| Target Modules | q\\\\\\\_proj, k\\\\\\\_proj, v\\\\\\\_proj, o\\\\\\\_proj, gate\\\\\\\_proj, up\\\\\\\_proj, down\\\\\\\_proj |
| Trainable Embeddings | embed\\\\\\\_tokens + lm\\\\\\\_head (8 yeni Lazca token eklendi) |
| Yeni Tokenler / New Tokens | Ç̌/ç̌, Ǩ/ǩ, P̌/p̌, Ť/ť, Ž/ž, Ʒ/ʒ, Ǯ/ǯ (Lazoğlu-Feurstein Latin alphabet) |
| Embedding Başlatma | phonetic_init — her token, Latin anahtar harf + Kartvelian/Gürcü harfinin embedding ortalaması ile başlar |
| Eğitim Verisi / Training Data | ~263.600 cümle çifti (laz↔tr: 87.870, laz↔en: 87.835) |
| Epoch / Epochs | 3 |
| Optimizer | AdamW, learning rate 1e-4, cosine schedule, warmup ratio 0.1 |
| Max Sequence Length | 512 tokens |
| Attention Backend | SDPA (cuDNN-backed) |
| Precision | bf16 |
| Gradient Checkpointing | Enabled (use\\\\\\\_reentrant=false) |
| Training Loss | 8.67 → 0.059 (step 10 → step 3090) |
| Token Accuracy | %98.5 (son adımlar) |
| Hardware | NVIDIA ( GB10 ) |
---
| Yön / Direction | Kaynak Diller / Source Languages |
|---|---|
| Türkçe → Lazca | tr→laz |
| İngilizce → Lazca | en→laz |
| İspanyolca → Lazca | es→laz |
| Almanca → Lazca | de→laz |
| Rusça → Lazca | ru→laz |
| Arapça → Lazca | ar→laz |
---
Aşağıdaki çıktılar, checkpoint-3090 modelinden greedy decoding ile alınmıştır. Model \\\*\\\*deneysel\\\*\\\* aşamadadır; sonuçlar doğrultuda ve içeriğe göre değişebilir.
The following outputs were generated from checkpoint-3090 via greedy decoding (temperature=0). This model is still experimental — results may vary by context and content.
| # | Türkçe (Kaynak) | Lazca (Çeviri) |
|---|---|---|
| 1 | Merhaba, nasılsın? | Ç'o, muç'o-ti si? |
| 2 | Bugün hava çok güzel. | Andğa ťaroni ar ǩelendo ǩaiyu. |
| 3 | Denize bugün gidebilir miyiz? | Andğa zuğaşa gamamalen-i? |
| 4 | Çay demlemek istiyorum. | Çayi dolobdvare. |
| 5 | Bu akşam ne yapacağız? | Amseri mu p̌aten? |
| # | English (Source) | Lazca (Translation) |
|---|---|---|
| 1 | Hello, how are you? | Ç'o, muç'o-ti si? |
| 2 | The weather is beautiful today. | Andğa mapxa on. |
| 3 | Can we go to the sea today? | Andğa zuğaşa ilen-i? |
| 4 | I want to make tea. | Çayi dolobobonas. |
| 5 | What are we doing tonight? | Amseri mu p̌aten? |
| # | Español (Fuente) | Lazca (Traducción) |
|---|---|---|
| 1 | Hola, ¿cómo estás? | Ç'o, mu xali? |
| 2 | Hoy hace mucho sol. | Andğa dido mjora ren. |
| 3 | ¿Podemos ir a la playa hoy? | Andğa noğaşa amaxta-i? |
| 4 | Quiero tomar un té. | Ar nçayi opşvana unon. |
| 5 | ¿Qué hacemos esta noche? | Amseri mu p̌aten? |
| # | Deutsch (Quelle) | Lazca (Übersetzung) |
|---|---|---|
| 1 | Hallo, wie geht es dir? | Ç'o, mu xali kogiç'o? |
| 2 | Das Wetter ist heute schön. | Andğa mapxa on. |
| 3 | Können wir heute ans Meer gehen? | Andğa noğaşe mzuğaşa moy-gulurt? |
| 4 | Ich möchte Tee trinken. | Çayi opşvana momalen. |
| 5 | Was machen wir heute Abend? | Amseri mu biç'opaten? |
| # | Русский (Источник) | Lazca (Перевод) |
|---|---|---|
| 1 | Привет, как дела? | Ç'o, mu xali? |
| 2 | Сегодня хорошая погода. | Andğa ǩai ťaoni ren. |
| 3 | Может, пойдем к морю сегодня? | Andğa zuğaşa vidat-i? |
| 4 | Я хочу пить чай. | Çayi opşvana unon. |
| 5 | Что мы будем делать вечером? | Seri mu p̌aten? |
| # | العربية (المصدر) | Lazca (الترجمة) |
|---|---|---|
| 1 | مرحبا، كيف حالك؟ | Ç'e, mu xali? |
| 2 | الجو جميل اليوم. | Andğa ťaoni msǩva on. |
| 3 | هل نستطيع الذهاب إلى البحر اليوم؟ | Andğa zuğaşa gamamaxtasen-i? |
| 4 | أريد أن أشرب الشاي. | Çayi opşva. |
| 5 | ماذا سنفعل الليلة؟ | Amseri mu p̌aten? |
---
Model, Lazoğlu-Feurstein Latin alfabesindeki Lazca'ya özgü karakterleri doğru bir şekilde üretir:
Ç̌ / ç̌ — c-cedilla + combining caron (U+00C7 + U+030C)
Ǩ / ǩ — k + caron (U+01E8 / U+01E9)
P̌ / p̌ — p + combining caron (U+0050 + U+030C)
Ť / ť — t + caron (U+0164 / U+0165)
Ž / ž — z + caron (U+017D / U+017E)
Ʒ / ʒ — ezh (U+01B7 / U+0292)
Ǯ / ǯ — ezh + caron (U+01EE / U+01EF)
Bu karakterler tokenizer'a normal kelime parçacıkları olarak eklenmiştir (special\\\\\\\_tokens=True değil) — bu sayede çeviri çıktısında kaybolmazlar.
The model correctly produces Lazca characters specific to the Lazoğlu-Feurstein Latin alphabet. These characters were added to the tokenizer as regular vocabulary tokens (not special\\\\\\\_tokens=True), so they are preserved in translation output rather than being stripped.
---
This model is experimental and trained on limited data (~263K sentence pairs). It works reasonably for everyday/practical sentences but has not been validated on literary or technical texts. Some outputs may be longer than necessary due to repeated generation patterns inherited from training format. This model is not available for download. It will be announced via anadil.org when it becomes available.
---
TranslateGemma'nın resmi chat template'i (ISO 639-1 dil kodları gerektirir) Lazca için uygun değildir — çünkü Lazca'nın bir ISO 639-1 kodu yoktur (yalnızca ISO 639-3: lzz). Bunun yerine, eğitim ve çıkarım Gemma'nın yerel turn token formatını kullanır:
<bos><start\\\\\\\_of\\\\\\\_turn>user\\\\\\\\nYou are a professional translation system.\\\\\\\\n\\\\\\\\nTranslate the following text from {kaynak} to {hedef}.\\\\\\\\n\\\\\\\\nTEXT:\\\\\\\\n{metin}<end\\\\\\\_of\\\\\\\_turn>\\\\\\\\n<start\\\\\\\_of\\\\\\\_turn>model\\\\\\\\n{cevap}<end\\\\\\\_of\\\\\\\_turn>\\\\\\\\n
Bu yaklaşım, modelin zaten güçlü bir önceliği olan <end\\\\\\\_of\\\\\\\_turn> durdurma token'ını kullanır ve LoRA ile yalnızca attention/projeksiyon katmanlarına etki eder. Lazca karakterleri için embedding'ler ise phonetic initialization yöntemiyle (Latin anahtar harf + Kartvelian/Gürcü harfinin embedding ortalaması) başlatılmıştır.
The official TranslateGemma chat template requires ISO 639-1 language codes, which doesn't support Lazca (only ISO 639-3: lzz). Instead, training and inference use Gemma's native turn token format, leveraging the model's strong pre-trained prior on <end\\\\\\\_of\\\\\\\_turn> as a stop signal. LoRA only affects attention/projection layers, and new Lazca token embeddings are initialized via phonetic initialization — averaging the embeddings of a Latin anchor letter and its Kartvelian/G Georgian counterpart.
---
Geri bildirim, hata raporları ve model güncellemeleri için:
For feedback, bug reports, and model updates: