Downloads · 30 days
6
15% of all-time downloads
WishAshake/XLM-Roberta
XLM-Roberta is a text classification model from WishAshake. Use it when you need a label for a piece of text. The card lists the license as mit.
A fine-tuned XLM-RoBERTa model for detecting hate speech and offensive content in Roman Urdu text.
Downloads · 30 days
6
15% of all-time downloads
All-time downloads
41
Public
Parameters
278M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
A fine-tuned XLM-RoBERTa model for detecting hate speech and offensive content in Roman Urdu text.
This model is based on xlm-roberta-base and has been fine-tuned on the Hate Speech Roman Urdu (HS-RU-20) dataset for binary classification:
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="WishAshake/XLM-Roberta"
)
# Classify text
result = classifier("your roman urdu text here")
print(result)
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
tokenizer = AutoTokenizer.from_pretrained("WishAshake/XLM-Roberta")
model = AutoModelForSequenceClassification.from_pretrained("WishAshake/XLM-Roberta")
# Tokenize and predict
text = "your roman urdu text here"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=128)
with torch.no_grad():
outputs = model(**inputs)
predictions = torch.nn.functional.softmax(outputs.logits, dim=-1)
label = "Toxic" if predictions[0][1] > 0.5 else "Safe"
confidence = predictions[0][1].item() if predictions[0][1] > 0.5 else predictions[0][0].item()
print(f"Label: {label}, Confidence: {confidence:.4f}")
The model was trained on the Hate Speech Roman Urdu (HS-RU-20) dataset, which contains:
If you use this model in your research, please cite:
@misc{xlm-roberta-roman-urdu-hate-speech,
title={XLM-RoBERTa for Roman Urdu Hate Speech Detection},
author={Wisha Zahid},
year={2024},
howpublished={\url{https://huggingface.co/WishAshake/XLM-Roberta}}
}
This model is released under the MIT License.
For questions or issues, please open an issue on the GitHub repository.