Downloads ยท 30 days
0
nahiar/twitter-bot-detection
twitter-bot-detection is a machine learning model from nahiar. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for scikit-learn. The card lists the license as apache-2.0.
This directory contains a trained Random Forest classifier for detecting bot accounts on Twitter.
Downloads ยท 30 days
0
Access
Public
Updated Jan 6, 2026
Repo size
222 MB
Likes
0
Public
Click a slice to open those files.
.pkl220 MB ยท 99%
From the Hugging Face model README
This directory contains a trained Random Forest classifier for detecting bot accounts on Twitter.
Model Version: v2 Training Date: 2025-11-27 12:08:54 Framework: scikit-learn 1.5.2 Algorithm: Random Forest Classifier with GridSearchCV Hyperparameter Tuning
| Metric | Score |
|---|---|
| Accuracy | 0.8771 (87.71%) |
| Precision | 0.8595 (85.95%) |
| Recall | 0.7558 (75.58%) |
| F1-Score | 0.8043 (80.43%) |
| ROC-AUC | 0.9354 (93.54%) |
| Average Precision | 0.9008 (90.08%) |
| File | Description |
|---|---|
twitter_bot_detection_v2.pkl | Trained Random Forest model |
twitter_scaler_v2.pkl | MinMaxScaler for feature normalization |
twitter_features_v2.json | List of features used by the model |
twitter_metrics_v2.txt | Detailed performance metrics report |
images/ | All visualization plots (13 images) |
README.md | This file |
Training Set:
Test Set:
has_custom_cover_imagedescription_lengthfavourites_countfollowers_countfriends_countfollowers_to_friends_ratiohas_locationusername_digit_countusername_lengthstatuses_countis_verifiedaccount_age_daysTotal combinations tested: 540
All visualizations are saved in the images/ directory:
import joblib
import pandas as pd
import numpy as np
# Load model and scaler
model = joblib.load('twitter_bot_detection_v2.pkl')
scaler = joblib.load('twitter_scaler_v2.pkl')
# Prepare your data (example)
data = {
'has_custom_cover_image': 0.5,
'description_length': 0.5,
'favourites_count': 0.5,
'followers_count': 0.5,
'friends_count': 0.5,
'followers_to_friends_ratio': 0.5,
'has_location': 0.5,
'username_digit_count': 0.5,
'username_length': 0.5,
'statuses_count': 0.5,
'is_verified': 0.5,
'account_age_days': 0.5,
}
# Create DataFrame
df = pd.DataFrame([data])
# Scale features
df_scaled = scaler.transform(df)
# Predict
prediction = model.predict(df_scaled)[0]
probability = model.predict_proba(df_scaled)[0]
print(f"Prediction: {'Bot' if prediction == 1 else 'Human'}")
print(f"Bot Probability: {probability[1]:.4f}")
print(f"Human Probability: {probability[0]:.4f}")
Predicted
Human Bot
Actual Human 4676 309
Bot 611 1891
To retrain the model:
../data/train_twitter.csv5_enhanced_training.ipynbFor questions or issues regarding this model, please refer to the main project documentation.
Generated: 2025-11-27 12:08:54
Notebook: 5_enhanced_training.ipynb
Platform: Twitter