Downloads · 30 days
26
6% of all-time downloads
chaichy/gbert-CTA-w-synth
gbert-CTA-w-synth is a text classification model from chaichy. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
gbert-CTA-w-synth is a fine-tuned version of the German BERT model (GBERT) designed to detect Calls to Action (CTAs) in political Instagram content. It was developed to analyze political mobilization strategies during…
Downloads · 30 days
26
6% of all-time downloads
All-time downloads
435
Public
Parameters
336M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
gbert-CTA-w-synth is a fine-tuned version of the German BERT model (GBERT) designed to detect Calls to Action (CTAs) in political Instagram content. It was developed to analyze political mobilization strategies during the 2021 German Federal Election, focusing on Instagram stories and posts.
This model is trained on real-world and synthetic data to mitigate class imbalances and improve performance. It specializes in detecting explicit and implicit CTAs in multimodal content, including captions, Optical Character Recognition (OCR) text from images, and video transcriptions.
deepset/gbert-largeFor video transcriptions, we used bofenghuang/whisper-large-v2-cv11-german, a fine-tuned version of OpenAI's Whisper model adapted for the German language.
The model was evaluated against human-annotated ground truth labels to ensure classification quality. We performed an evaluation using five-fold cross-validation to validate the model’s generalizability. The model was benchmarked with the following metrics:
The evaluation was based on a dataset containing 1,388 documents annotated by nine contributors. Disagreements were resolved using majority decisions.
This model is intended for computational social science and political communication research, specifically for studying how political actors mobilize audiences on social media. It is effective for detecting Calls to Action in German-language social media content.
You can use this model with the transformers library in Python:
from transformers import BertTokenizer, BertForSequenceClassification
import torch
# Load model and tokenizer
tokenizer = BertTokenizer.from_pretrained('chaichy/gbert-CTA-w-synth')
model = BertForSequenceClassification.from_pretrained('chaichy/gbert-CTA-w-synth')
# Tokenize input
inputs = tokenizer("Input text here", return_tensors="pt")
# Get classification results
outputs = model(**inputs)
logits = outputs.logits
predicted_class = torch.argmax(logits, dim=1)
# 0 for absence, 1 for presence of CTA
print(f"Predicted class: {predicted_class.item()}")
The model was trained on Instagram content collected during the 2021 German Federal Election campaign. This included:
The dataset contains both explicit and implicit CTAs, which are binary labeled (True/False). We generated synthetic training data based on the original human-annotated dataset to handle class imbalance. The synthetic dataset was created using OpenAI’s GPT-4o, which mimicked real-world CTAs by generating new examples in a consistent political communication style.
The training data was collected from publicly available Instagram posts and stories shared by verified political accounts during the 2021 German Federal Election. No personal or sensitive data was included.
If you use this model, please cite the following:
@misc{achmanndenkler2024detectingcallsactionmultimodal,
title={Detecting Calls to Action in Multimodal Content: Analysis of the 2021 German Federal Election Campaign on Instagram},
author={Michael Achmann-Denkler and Jakob Fehle and Mario Haim and Christian Wolff},
year={2024},
eprint={2409.02690},
archivePrefix={arXiv},
primaryClass={cs.SI},
url={https://arxiv.org/abs/2409.02690},
}