Downloads · 30 days
12
28% of all-time downloads
raviearjun/engagement-predictor
engagement-predictor is a tabular regression model from raviearjun. Use it for the tabular regression task on the model card, and read the license before you ship it in a product. It is set up for skops. The card lists the license as mit.
A regression model (scikit-learn Ridge inside a Pipeline) that predicts engagementscore for brand-creator campaign content, part of the Viralst application.
Downloads · 30 days
12
28% of all-time downloads
All-time downloads
43
Public
Repo size
231 KB
Likes
0
Public
Click a slice to open those files.
.csv562 KB · 70%
From the Hugging Face model README
A regression model (scikit-learn Ridge inside a Pipeline) that predicts
engagement_score for brand-creator campaign content, part of the Viralst
application.
Given campaign brief attributes (tone, hook type, target audience, creator
tier, objective, etc. — typically extracted from a brief document via an LLM
parser) and planned content execution details (duration, upload time,
caption), the model predicts the expected engagement_score
((likes + comments + shares) / views).
It is served through the POST /engagement/predict endpoint of the Viralst
ai/ service.
The model is trained on a synthetically generated dataset (2,000 rows,
300 unique briefs, 30 unique brands) produced by
training/generate_poc_data.py. Feature-to-target relationships were
deliberately encoded (with variable noise per row) so the training pipeline
— feature engineering, model selection, evaluation, and serving — could be
validated end-to-end ahead of onboarding real campaign data. Results below
characterize how well the model recovers the encoded synthetic relationships,
not real-world campaign performance.
sklearn.linear_model.Ridge inside a Pipeline (OneHotEncoder
for single-value categorical features, StandardScaler for numeric and
multi-hot features).GroupKFold (grouped by brief_id),
so evaluation reflects generalization to unseen briefs rather than
leaking the same brief across train/test folds.skops.io.dump (not raw pickle) for safer loading.{
"alpha": 10.0,
"n_rows": 2000,
"n_groups": 300,
"cv_mae": 0.011419824566734465,
"cv_rmse": 0.01458130687031906,
"cv_r2": 0.7102257580449983,
"baseline_mae": 0.021489796999999998,
"baseline_rmse": 0.027087352583401354
}
Mean-predictor baseline RMSE: 0.027087352583401354. The final model outperforms
this baseline by a wide margin on the synthetic dataset described above.
See config.json in this repo for the full list and order of feature
columns the model expects. The canonical definition lives in
src/engagement_predictor/features.py in the application repository
(https://github.com/<your-org>/viralst).
The raw synthetic dataset used to train this model is included in this repo
under dataset/ (brands.csv, briefs.csv, contents.csv — 30
brands, 300 briefs, 2000 content rows). It was generated by
training/generate_poc_data.py; see "Training data and methodology" above
for why it is synthetic and what it does and does not represent.
import pandas as pd
import skops.io as sio
from huggingface_hub import hf_hub_download
path = hf_hub_download(repo_id="raviearjun/engagement-predictor", filename="engagement_ridge.skops")
untrusted = sio.get_untrusted_types(file=path)
model = sio.load(path, trusted=untrusted)
# X must be a DataFrame with columns matching config.json -> feature_columns
prediction = model.predict(X)
Training environment (see requirements.txt in this repo for exact pins):
scikit-learn==1.7.2
skops==0.14.0
numpy==1.26.4
pandas==2.3.3