Downloads · 30 days
0
AzeerDev/LFM2.5-1.2B-Instruct-Saudi-Dialect
LFM2.5-1.2B-Instruct-Saudi-Dialect is a machine learning model from AzeerDev. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This model is a fine-tuned version of Liquid AI’s LFM2.5‑1.2B‑Instruct, adapted for Saudi dialect conversational generation.
Downloads · 30 days
0
Access
Public
Updated Aug 25, 2026
Repo size
14.2 MB
Likes
0
Public
Click a slice to open those files.
.safetensors14.2 MB · 75%
From the Hugging Face model README
This model is a fine-tuned version of Liquid AI’s LFM2.5‑1.2B‑Instruct, adapted for Saudi dialect conversational generation.
The base model belongs to the LFM2.5 family — hybrid state-space + attention language models designed for fast on-device inference,low memory usage, and strong performance relative to size. It contains ~1.17B parameters, 32k context length, and supports multilingual generation including Arabic.
This fine-tuned variant specializes the model for Saudi dialect conversational patterns, improving fluency, dialect authenticity, and instruction following for regional Arabic use cases.
Fine-tuned on:
Dataset:
HeshamHaroon/saudi-dialect-conversations
Domain: Conversational dialogue
Language: Saudi dialect Arabic
Format: Instruction → Response pairs
Purpose: Increase dialect authenticity and conversational naturalness.
(Extracted from training notebook)
| Parameter | Value |
|---|---|
| Epochs | 4 |
| Learning Rate | 2e-4 |
| Batch Size | 16 |
| Gradient Accumulation | 4 |
| Optimizer | AdamW |
| LR Scheduler | Linear |
| Warmup Ratio | 0.03 |
| Sequence Length | 8096 |
| Precision | FP16 |
| Training Type | Supervised Fine-Tuning (SFT) |
Training was performed using:
The base model weights were adapted rather than retrained from scratch.
Qualitative evaluation indicates:
Dialect-specific fine-tuning is known to significantly increase dialect generation accuracy and reduce standard-Arabic drift in Arabic LLMs.
Strengths
Limitations
Potential risks:
Mitigations:
Runs efficiently on:
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "AyoubChLin/lfm2.5-saudi-dialect"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
prompt = "تكلم باللهجة السعودية عن القهوة"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Same as base model license unless otherwise specified.
If you use this model:
@misc{saudi-dialect-lfm2.5,
author = {Cherguelaine Ayoub},
title = {Saudi Dialect LFM2.5},
year = {2026},
publisher = {Hugging Face}
}