Downloads · 30 days
56
100% of all-time downloads
LousisOfficial/LousisAI-4B
LousisAI-4B is a machine learning model from LousisOfficial. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Lousis, Qwen/Qwen3-4B-Base temel alınarak, ~7,810 satırlık bir Türkçe veri seti üzerinde QLoRA (4-bit taban model + LoRA adapter) ile eğitilmiş bir causal language modeldir. Bu bir sıfırdan eğitim değil, hazır bir pre…
Downloads · 30 days
56
100% of all-time downloads
All-time downloads
56
Public
Parameters
4B
3.4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors3.4 GB · 100%
How the weights are stored.
U83.6B · 90%
From the Hugging Face model README
Lousis, Qwen/Qwen3-4B-Base temel alınarak, ~7,810 satırlık bir Türkçe
veri seti üzerinde QLoRA (4-bit taban model + LoRA adapter)
ile eğitilmiş bir causal language modeldir.
Bu bir sıfırdan eğitim değil, hazır bir pretrained model üzerine yapılan adaptasyondur.
Taban model yaklaşık 2238.8M parametre. Bunların yaklaşık %1.48'i (LoRA adapter, r=16) eğitilmiştir; kaydedilen model, adapter taban modelle birleştirilmiş halidir.
Qwen/Qwen3-4B-Base ile birlikte gelen orijinal tokenizer kullanılmıştır (kendi
tokenizer'ımız eğitilmemiştir).
Google Colab, NVIDIA Tesla T4 (~15 GB VRAM), fp16 hesaplama, 8-bit AdamW, gradient checkpointing, gradient accumulation, 4-bit NF4 quantization (QLoRA).
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("LousisOfficial/LousisAI-4B")
model = AutoModelForCausalLM.from_pretrained("LousisOfficial/LousisAI-4B")
# Model ChatML formatiyla egitildi (<|im_start|> / <|im_end|>)
prompt = "<|im_start|>user\nMerhaba<|im_end|>\n<|im_start|>assistant\n"
inputs = tokenizer(prompt, return_tensors="pt")
im_end = tokenizer.convert_tokens_to_ids("<|im_end|>")
outputs = model.generate(**inputs, max_new_tokens=100, do_sample=True, temperature=0.5,
eos_token_id=[im_end, tokenizer.eos_token_id])
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Bu model Qwen/Qwen3-4B-Base temel alınarak fine-tune edilmiştir; sıfırdan
eğitilmiş bir model DEĞİLDİR.