Downloads · 30 days
24
26% of all-time downloads
ringorsolya/guilt-classifier-en
guilt-classifier-en is a machine learning model from ringorsolya. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-4.0.
GuiltRoBERTa-en is a two-stage AI pipeline for detecting guilt-assignment rhetoric in English political discourse. It combines:
Downloads · 30 days
24
26% of all-time downloads
All-time downloads
93
Public
Parameters
278M
1.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors1.1 GB · 100%
From the Hugging Face model README
GuiltRoBERTa-en is a two-stage AI pipeline for detecting guilt-assignment rhetoric in English political discourse. It combines:
guilt vs no_guilt)The approach is grounded in political communication theory, which suggests that guilt attribution often emerges in anger-laden contexts. Thus, only texts labeled as "Anger" in Stage 1 are passed to the guilt classifier.
Anger, Fear, Disgust, Sadness, Joy, None of them)predicted_emotion == "Anger" for Stage 2The Babel Emotions Tool is not an API but a web-based interface. Upload a CSV file, download the labeled results, and use them as input to the guilt classifier.
xlm-roberta-baseguilt, no_guilt)Guilt assignment — attributing moral responsibility or blame — is a key rhetorical strategy in political communication. Since guilt often appears alongside anger, direct one-stage classification risks conflating emotional tones.
This two-stage pipeline improves precision by:
The model was evaluated on a held-out validation set (20% stratified split) with the following approach:
| Stage 1 Filter | Threshold (τ) | Precision | Recall | F1 | Accuracy |
|---|---|---|---|---|---|
| Anger-only | 0.15 | optimized | optimized | optimized | optimized |
emotion_predicted column)import pandas as pd
from transformers import AutoTokenizer, AutoModelForSequenceClassification, TextClassificationPipeline
# Load Babel emotion predictions
df = pd.read_excel("your_data_with_emotion_predictions.xlsx")
# Filter for 'Anger' only
anger_df = df[df["emotion_predicted"] == "Anger"].copy()
# Load the guilt classifier
repo_id = "ringorsolya/guilt-classifier-en"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSequenceClassification.from_pretrained(repo_id)
pipe = TextClassificationPipeline(model=model, tokenizer=tokenizer, return_all_scores=True)
# Apply guilt predictions with threshold
THRESHOLD = 0.15
anger_df["guilt_score"] = anger_df["text"].apply(
lambda t: pipe(t)[0][1]["score"] # score for 'guilt' label
)
anger_df["guilt_predicted"] = anger_df["guilt_score"] > THRESHOLD
# Save results
anger_df.to_excel("anger_with_guilt_predictions.xlsx", index=False)
# Statistics
print(f"Total anger sentences: {len(anger_df)}")
print(f"Predicted guilt: {anger_df['guilt_predicted'].sum()}")
print(f"Guilt ratio: {anger_df['guilt_predicted'].mean():.2%}")
import torch
from transformers import XLMRobertaTokenizer, XLMRobertaForSequenceClassification
# Load model
model_path = "your-org/guiltroberta-en"
tokenizer = XLMRobertaTokenizer.from_pretrained(model_path)
model = XLMRobertaForSequenceClassification.from_pretrained(model_path)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)
model.eval()
# Example: anger-labeled sentence
text = "I'm furious at myself for letting this happen again."
# Tokenize and predict
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512, padding=True)
inputs = {k: v.to(device) for k, v in inputs.items()}
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits
prob_guilt = torch.softmax(logits, dim=-1)[0][1].item()
# Apply threshold
THRESHOLD = 0.15
prediction = "guilt" if prob_guilt > THRESHOLD else "no_guilt"
print(f"Guilt probability: {prob_guilt:.4f}")
print(f"Prediction: {prediction}")
Epochs: 4
Learning Rate: 2e-5
Batch Size: 8
Max Sequence Length: 512 tokens
Optimizer: AdamW
Scheduler: Linear warmup
Train/Validation Split: 80/20 (stratified)
Class Weighting: Applied to handle label imbalance