Downloads · 30 days
11
20% of all-time downloads
yrrhall/dialect-router-v0.1
dialect-router-v0.1 is a text classification model from yrrhall. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as mit.
A lightweight Arabic dialect identification model that classifies input text into one of 11 Arabic dialect / language codes. It is used as the routing backbone in the Lahgtna pipeline to automatically select the corre…
Downloads · 30 days
11
20% of all-time downloads
All-time downloads
56
Public
Parameters
11.6M
46.2 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors46.2 MB · 97%
From the Hugging Face model README
A lightweight Arabic dialect identification model that classifies input text into one of 11 Arabic dialect / language codes. It is used as the routing backbone in the Lahgtna pipeline to automatically select the correct voice reference and Chatterbox language token for speech synthesis.
| Property | Value |
|---|---|
| Architecture | Transformer encoder (sequence classification head) |
| Task | Multi-class text classification |
| Input | Raw Arabic text (up to 512 tokens) |
| Output | One of 11 dialect codes |
| Language | Arabic (ar) |
| License | MIT |
| Label | Dialect | Region |
|---|---|---|
eg | Egyptian | Egypt |
sa | Saudi | Saudi Arabia |
mo | Moroccan (Darija) | Morocco |
iq | Iraqi | Iraq |
sd | Sudanese | Sudan |
tn | Tunisian | Tunisia |
lb | Lebanese | Lebanon |
sy | Syrian | Syria |
ly | Libyan | Libya |
ps | Palestinian | Palestine |
ar | Modern Standard Arabic (MSA) | — |
Dialect-aware TTS routing — given an Arabic utterance, predict the dialect so the correct speaker reference audio and Chatterbox language code can be selected automatically.
Standalone Arabic dialect identification for NLP pipelines, content filtering, dataset analysis, or any application that needs to distinguish Arabic dialects programmatically.
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_id = "oddadmix/dialect-router-v0.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()
text = "اه ياراسي الواحد دماغه وجعاه"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
logits = model(**inputs).logits
pred_id = torch.argmax(logits, dim=-1).item()
dialect = model.config.id2label[pred_id]
print(dialect) # e.g. "eg"
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="oddadmix/dialect-router-v0.1",
)
result = classifier("اه ياراسي الواحد دماغه وجعاه")
print(result)
# [{'label': 'eg', 'score': 0.94}]
from inference import run_pipeline
# Dialect is detected automatically
run_pipeline(
text="اه ياراسي الواحد دماغه وجعاه",
output_path="output.wav",
)
sy / lb, eg / ly) may be confused by the model.sd, ly) may have lower recall.If you use this model in your research or product, please cite:
@misc{lahgtna-dialect-router-2025,
title = {dialect-router-v0.1: Arabic Dialect Identification for TTS Routing},
author = {Oddadmix},
year = {2025},
url = {https://huggingface.co/oddadmix/dialect-router-v0.1}
}
oddadmix/lahgtna-chatterbox-v1