Downloads · 30 days
23
22% of all-time downloads
GhadeerALbadani/XLM_Latin
XLM_Latin is a fill-mask model from GhadeerALbadani. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
XLMLatin is a transliteration-aware multilingual language model developed to investigate the impact of script unification on multilingual and cross-lingual hate speech detection. The model is based on the XLM-RoBERTa…
Downloads · 30 days
23
22% of all-time downloads
All-time downloads
103
Public
Parameters
111M
443 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors443 MB · 100%
From the Hugging Face model README
XLM_Latin is a transliteration-aware multilingual language model developed to investigate the impact of script unification on multilingual and cross-lingual hate speech detection. The model is based on the XLM-RoBERTa architecture and was further pretrained on multilingual transliterated corpora represented in a unified Latin script.
The model is intended for:
The model was trained on transliterated text from:
The model was pretrained on 360,768 transliterated text samples collected from multiple sources, including news articles, sentiment datasets, social media content, and public corpora.
The model was evaluated using:
The model processes hate speech data that may contain offensive content. It should not be used as the sole decision-making system in high-risk applications.
Install the required libraries:
pip install transformers torch sentencepiece
The pretrained XLM_Latin model can be loaded directly from the Hugging Face Hub:
from transformers import AutoTokenizer, AutoModelForMaskedLM
model_name = "GhadeerALbadani/XLM_Latin"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForMaskedLM.from_pretrained(model_name)
Replace YOUR_USERNAME with your Hugging Face username or organization name.
The tokenizer can be used to convert transliterated multilingual text into model inputs:
text = "ana la uhibbu hadha almujtama"
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
padding=True
)
print(inputs)
The model can be used to generate contextual embeddings for transliterated multilingual text:
from transformers import AutoTokenizer, AutoModel
model_name = "GhadeerALbadani/XLM_Latin"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)
text = "ana la uhibbu hadha almujtama"
inputs = tokenizer(
text,
return_tensors="pt"
)
outputs = model(**inputs)
embeddings = outputs.last_hidden_state
print(embeddings.shape)
Since XLM_Latin is pretrained using the Masked Language Modeling (MLM) objective, it can predict masked tokens:
from transformers import pipeline
fill_mask = pipeline(
"fill-mask",
model="GhadeerALbadani/XLM_Latin",
tokenizer="GhadeerALbadani/XLM_Latin"
)
result = fill_mask(
"ana <mask> hadha almujtama"
)
print(result)
XLM_Latin is intended to serve as a pretrained backbone model and can be further fine-tuned for multilingual hate speech detection tasks:
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained(
"YOUR_USERNAME/XLM_Latin",
num_labels=2
)
The resulting fine-tuned model can then be used for monolingual, multilingual, cross-lingual, and zero-shot hate speech detection experiments.
@mastersthesis{XLM_Latin_2026,
title={A Multilingual Model for Detecting Hate Speech on Unknown Language Script},
author={Author Name},
year={2026},
school={University Name}
}