Downloads · 30 days
18
40% of all-time downloads
Adnan855570/Roberta_base_model
Roberta_base_model is a text classification model from Adnan855570. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as other.
- Base model: urduhack/roberta-urdu-small - Task: Binary text classification (hate vs. nothate) - Language: Urdu (ur) - Labels - 0 → nothate - 1 → hate
Downloads · 30 days
18
40% of all-time downloads
All-time downloads
45
Public
Parameters
125M
499 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors499 MB · 99%
From the Hugging Face model README
urduhack/roberta-urdu-smallnot_hatehateThis model fine-tunes a small RoBERTa for Urdu hate-speech detection. Class imbalance was addressed by oversampling with SMOTE at the feature level (TF–IDF) prior to tokenization-based training.
Adnan855570/urdu-hate-speech (Excel files: preprocessed_combined_file (1).xlsx, Urdu_Hate_Speech.xlsx)Tweet (text), Tag (label in {0,1})AutoTokenizer.from_pretrained("urduhack/roberta-urdu-small") with truncation=True, padding=TrueAutoModelForSequenceClassification with num_labels=2Note: Results derive from the balanced (SMOTE) dataset and the 80/20 split used in the notebook.
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
MODEL_ID = "Adnan855570/urdu-roberta-hate" # replace if different
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_ID).eval()
id2label = model.config.id2label or {"0":"not_hate","1":"hate"}
def predict(text: str):
enc = tokenizer(text, return_tensors="pt", truncation=True, padding=True)
with torch.no_grad():
logits = model(**enc).logits
probs = logits.softmax(dim=-1).squeeze().tolist()
pred = int(logits.argmax(dim=-1).item())
return {"label_id": pred, "label": id2label.get(str(pred), str(pred)),
"scores": {"not_hate": probs[0], "hate": probs[1]}}
print(predict("یہ نفرت انگیز ہے یا نہیں؟"))
Or with a pipeline:
from transformers import pipeline
clf = pipeline("text-classification", model="Adnan855570/urdu-roberta-hate", top_k=None)
print(clf("یہ نفرت انگیز ہے یا نہیں؟"))
curl -X POST -H "Authorization: Bearer $HF_TOKEN" -H "Content-Type: application/json" \
-d '{"inputs":"یہ نفرت انگیز ہے یا نہیں؟"}' \
https://api-inference.huggingface.co/models/Adnan855570/urdu-roberta-hate
import os, requests
API_URL = "https://api-inference.huggingface.co/models/Adnan855570/urdu-roberta-hate"
HEADERS = {"Authorization": f"Bearer {os.environ.get('HF_TOKEN','')}"}
print(requests.post(API_URL, headers=HEADERS, json={"inputs":"..."}, timeout=30).json())
Ensure the config includes:
id2label = {"0":"not_hate","1":"hate"}label2id = {"not_hate":0,"hate":1}random_state=42urduhack/roberta-urdu-small). Set a compatible license here once confirmed.@misc{urdu_roberta_hate_balanced_2025,
title = {Urdu RoBERTa Hate Speech Classifier (Balanced)},
author = {Adnan},
year = {2025},
howpublished = {\url{https://huggingface.co/Adnan855570/urdu-roberta-hate}}
}
urduhack/roberta-urdu-small