Downloads · 30 days
462
29% of all-time downloads
thealper2/MiniCPM5-1B-Turkish
MiniCPM5-1B-Turkish is a text generation model from thealper2. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
openbmb/MiniCPM5-1B tabanlı, Türkçe talimat takibi ve kod üretimi için tam ince ayar (full fine-tuning) yapılmış bir sohbet modeli.
Downloads · 30 days
462
29% of all-time downloads
All-time downloads
1.6K
Public
Parameters
1.1B
4.8 GB on disk
Likes
3
Public
Click a slice to open those files.
.gguf2.6 GB · 55%
From the Hugging Face model README
openbmb/MiniCPM5-1B tabanlı, Türkçe talimat takibi ve kod üretimi için tam ince ayar (full fine-tuning) yapılmış bir sohbet modeli.
A fully fine-tuned (not LoRA/QLoRA) chat model derived from openbmb/MiniCPM5-1B, targeting Turkish instruction following and Turkish-instructed code generation.
| Taban model / Base model | openbmb/MiniCPM5-1B |
| Mimari / Architecture | LlamaForCausalLM (MiniCPM5) |
| Parametre / Parameters | 1,080,632,832 |
| Eğitilen parametre / Trained | 1,080,632,832 (100.0%) |
| Yöntem / Method | Supervised fine-tuning (SFT), full-parameter |
| Diller / Languages | Türkçe (birincil), İngilizce (korunmuş) |
| Bağlam / Context | eğitim 2048 token (taban model 131k) |
| Precision | torch.bfloat16 |
| Lisans / License | apache-2.0 |
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "thealper2/MiniCPM5-1B-Turkish"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, dtype=torch.bfloat16, device_map="auto"
)
messages = [
{"role": "user", "content": "Python'da bir CSV dosyasını okuyup eksik değerleri temizleyen bir fonksiyon yaz."}
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
# do_sample=True sarttir - asagidaki nota bakin / see the decoding note below
out = model.generate(inputs, max_new_tokens=512, do_sample=True, temperature=0.7, top_p=0.95)
print(tokenizer.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))
Sohbet şablonu taban modelin ChatML türevidir (
<|im_start|>role\n...<|im_end|>). Her zamanapply_chat_templatekullanın; prompt'u elle kurmayın.
Model, yanıtlarına boş bir düşünme bloğu (<think>\n\n</think>) ile başlayacak şekilde eğitildi; bu, taban modelin enable_thinking=False biçimiyle uyumludur. Uzun zincirleme akıl yürütme (CoT) verisiyle eğitilmedi.
llama-cli -hf thealper2/MiniCPM5-1B-Turkish -p "Merhaba, kendini tanıt."
# veya yerel dosyayla:
llama-cli -m MiniCPM5-1B-Turkish-Q8_0.gguf -cnv
| dosya / file | tür / type | boyut / size |
|---|---|---|
MiniCPM5-1B-Turkish-BF16.gguf | BF16 | 2066 MB |
MiniCPM5-1B-Turkish-Q4_K_M.gguf | Q4_K_M | 656 MB |
MiniCPM5-1B-Turkish-Q5_K_M.gguf | Q5_K_M | 750 MB |
MiniCPM5-1B-Turkish-Q8_0.gguf | Q8_0 | 1100 MB |
Tüm kaynaklar tek bir konuşma şemasına (messages) normalize edildi; ardından temizleme, tüm veri setleri arasında birebir tekrar temizliği (exact dedup) ve kategori bazlı ağırlıklı örnekleme uygulandı. Veri setleri körlemesine birleştirilmedi.
| kaynak / source | kategori | ham satır | temizlik sonrası | seçilen |
|---|---|---|---|---|
AlicanKiraz0/Turkce-Atlas-Instruct | general_turkish | 336,146 | 336,099 | 55,000 |
tascib/turkish-instruction | general_turkish | 324,080 | 311,627 | 55,000 |
sixfingerdev/turkish-qa-multi-dialog-dataset | qa_dialog | 21,282 | 18,790 | 20,000 |
berhaan/Turkish-CodeAlpaca-20k | coding | 19,996 | 19,565 | 19,474 |
alztrk/turkish-code-instructions | coding | 2,676 | 1,895 | 1,895 |
bysismo/Turkish-Python-instruction-500k | coding | 335,286 | 330,668 | 48,631 |
| ayar / setting | değer / value |
|---|---|
| Yöntem | Full fine-tuning (LoRA/QLoRA kullanılmadı) |
| Kayıp / Loss | Yalnızca asistan token'ları (kullanıcı ve sistem token'ları -100 ile maskelendi) |
| Epoch | 1 |
| Learning rate | 2e-05 |
| Scheduler | cosine (warmup 0.03) |
| Optimizer | paged_adamw_8bit |
| Weight decay | 0.1 |
| Grad clipping | 1.0 |
| Batch | 1 x 16 accum |
| Sekans uzunluğu | 2048 |
| Sequence packing | True |
| Gradient checkpointing | True |
| Precision | torch.bfloat16 |
| Donanım / Hardware | NVIDIA GeForce RTX 5060 Ti |
2.04832324981689461.30104722976684561.2769154310226445.435 GBDeğerlendirme, taban model ile ince ayarlı model üzerinde aynı prompt'lar ve aynı çözümleme ayarlarıyla yapıldı. Türkçe kazanımının genel bir iyileşme anlamına gelmediğini görebilmek için İngilizce ve akıl yürütme yetenekleri de ayrıca ölçüldü (regresyon testi).
| sonuç / verdict | prompt |
|---|---|
| improved | 9 |
| not auto-scored | 9 |
| possible regression | 2 |
| unchanged | 10 |
Olası gerilemeler / possible regressions
code_js_01 (code_generation): 0/1 vs 1/1 checksen_code_02 (english_coding): 0/2 vs 2/2 checksalztrk/turkish-code-instructions) Türkçe karakterleri ASCII'ye indirgenmiş metin içerir; bu nedenle katkısı bilinçli olarak düşük tutuldu.python scripts/inspect_datasets.py
python scripts/prepare_datasets.py
python scripts/train.py --dry_run
python scripts/train.py
python scripts/compare_models.py --run
Tüm hiperparametreler config/training.yaml içindedir ve eğitim çıktısıyla birlikte config.yaml, dataset_report.json, training_summary.json olarak kaydedilir.
@misc{minicpm5_turkish,
title = {MiniCPM5-1B-Turkish},
note = {Full supervised fine-tune of openbmb/MiniCPM5-1B for Turkish instruction following and coding},
year = {2026}
}
Taban model / base model: openbmb/MiniCPM5-1B