Downloads · 30 days
0
kohils/Twitter-Cyberbullying-Classification
Twitter-Cyberbullying-Classification is a text classification model from kohils. Use it when you need a label for a piece of text. It is set up for sklearn.
This is a traditional Machine Learning model that classifies tweets into different categories of cyberbullying. It is an ensemble Voting Classifier combining Logistic Regression and Random Forest, achieving approximat…
Downloads · 30 days
0
Access
Public
Updated Jan 29, 2026
Repo size
179 MB
Likes
1
Public
Click a slice to open those files.
.pickle93.6 MB · 100%
From the Hugging Face model README
This is a traditional Machine Learning model that classifies tweets into different categories of cyberbullying. It is an ensemble Voting Classifier combining Logistic Regression and Random Forest, achieving approximately 91% accuracy.
This model is designed to detect specific types of cyberbullying in text. It is lightweight and faster than transformer models, making it suitable for low-resource environments.
The model classifies text into 5 categories (mapped as follows):
0: Not Cyberbullying1: Gender (Sexist)2: Religion3: Age4: Ethnicity (Racist)(Note: The 'Other' category was removed during preprocessing to improve accuracy.)
The text underwent rigorous cleaning using the tweet-preprocessor library and custom functions:
WordNetLemmatizer.LogisticRegression (C=100, penalty='l2')RandomForestClassifier (n_estimators=100)VotingClassifier (Hard Voting) combining the above two.To use this model in Python, you need to load both the vectorizer and the model using joblib.
import joblib
import preprocessor as p # pip install tweet-preprocessor
import string
# 1. Load the saved files
model = joblib.load('model.pickle')
vectorizer = joblib.load('tfidf.pickle')
# 2. Define the cleaning function (Must match training!)
def clean_text(text):
text = p.clean(text)
text = text.lower()
text = "".join([char for char in text if char not in string.punctuation])
return text
# 3. Make a prediction
text = "You are dumb and you should go back to school."
clean_input = clean_text(text)
# Vectorize the text
vectorized_input = vectorizer.transform([clean_input])
# Predict
prediction = model.predict(vectorized_input)
classes = {0: 'Not Cyberbullying', 1: 'Gender', 2: 'Religion', 3: 'Age', 4: 'Ethnicity'}
print(f"Prediction: {classes[prediction[0]]}")