Downloads · 30 days
25
32% of all-time downloads
k12tr/mini-tr-9m
mini-tr-9m is a text generation model from k12tr. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Llama-TR-Mini is an experimental, ultra-lightweight Turkish language model with 9.3 million parameters, trained from scratch using the Llama 3 architecture.
Downloads · 30 days
25
32% of all-time downloads
All-time downloads
78
Public
Parameters
9.3M
37.4 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors37.4 MB · 98%
From the Hugging Face model README
Llama-TR-Mini is an experimental, ultra-lightweight Turkish language model with 9.3 million parameters, trained from scratch using the Llama 3 architecture.
This project was developed to explore the limits of small-scale language modeling and to understand the end-to-end pre-training/fine-tuning pipeline on consumer-grade hardware (Apple Silicon).
The model was trained on the Turkish-Alpaca dataset, which contains approximately 52K instruction-following pairs translated into Turkish.
Important Note: Due to its extremely small size (9M parameters), this model is prone to significant hallucinations and may produce nonsensical or repetitive outputs.
<|start_header_id|>user<|end_header_id|>).You can load this model using the transformers library:
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "k12tr/mini-tr-9m"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
prompt = "<|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\nTürkiye'nin başkenti neresidir?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n"
inputs = tokenizer(prompt, return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=50, temperature=0.1, repetition_penalty=1.5)
print(tokenizer.decode(output[0], skip_special_tokens=True))