Downloads · 30 days
0
KABURAKURIA/lugha1
lugha1 is a machine learning model from KABURAKURIA. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
lugha1 is an instruction-tuned and preference-aligned conversational language model derived from SmolLM-1.7B-Instruct. It is optimized for clear instruction following, polite dialogue, and socially grounded responses,…
Downloads · 30 days
0
Access
Public
Updated Jan 5, 2026
Repo size
3.2 MB
Likes
0
Public
Click a slice to open those files.
.json4.3 MB · 54%
From the Hugging Face model README
lugha1 is an instruction-tuned and preference-aligned conversational language model derived from SmolLM-1.7B-Instruct. It is optimized for clear instruction following, polite dialogue, and socially grounded responses, with an emphasis on small-talk, explanations, and helpful conversational tone.
The model was trained using Supervised Fine-Tuning (SFT) and further aligned using Direct Preference Optimization (DPO) with parameter-efficient fine-tuning (LoRA).
Base model: HuggingFaceTB/SmolLM-1.7B-Instruct
Architecture: Decoder-only Transformer
Parameters: ~1.7B (LoRA-adapted)
Fine-tuning method:
Precision: 4-bit (training), fp16 compute
Language(s): English (primary); conversational multilingual robustness inherited from base model
Model name: lugha1
The model was first trained using supervised instruction–response pairs, formatted in a conversational style:
### Instruction:
<user instruction>
### Response:
<assistant response>
This stage teaches the model:
After SFT, the model was aligned using Direct Preference Optimization (DPO). Preference pairs (chosen vs. rejected responses) were used to bias the model toward:
DPO was chosen over RLHF to ensure:
Training data consisted of instruction-style conversational datasets inspired by Smalltalk-style interactions, including:
Note: Some preference data was synthetically generated for alignment purposes.
lugha1 is suitable for:
Example use cases:
This model is not intended for:
Users should apply appropriate validation and human oversight.
from transformers import pipeline
pipe = pipeline(
"text-generation",
model="your-username/lugha1",
tokenizer="your-username/lugha1",
device_map="auto"
)
prompt = """### Instruction:
Explain empathy using humor.
### Response:"""
pipe(prompt, max_new_tokens=150)
Platform: Google Colab (free tier)
GPU: NVIDIA T4 / L4
Frameworks:
This model inherits the license of its base model: SmolLM-1.7B-Instruct license (see base model card for details).