Downloads · 30 days
14
23% of all-time downloads
mmochtak/lieline
lieline is a machine learning model from mmochtak. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-4.0.
LieLine is a fine-tuned RoBERTa-base classifier that detects disinformation / lie allegations — instances where a political speaker explicitly or implicitly accuses another actor of lying, deception, or spreading disi…
Downloads · 30 days
14
23% of all-time downloads
All-time downloads
60
Public
Parameters
125M
499 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors499 MB · 99%
From the Hugging Face model README
LieLine is a fine-tuned RoBERTa-base classifier that detects disinformation / lie allegations — instances where a political speaker explicitly or implicitly accuses another actor of lying, deception, or spreading disinformation — in political speech text. It was developed as part of the pipeline described in "Finding the Needle in a Haystack: Using Large Language Models to Detect Rare Speech Acts" (Mochtak & Meijers, 2026).
roberta-base (English)1/0)simpletransformersLieLine is intended for research use in computational social science and political communication research, specifically for:
The training data was constructed through a multi-phase pipeline rather than through direct random sampling of the underlying corpus, to address the rarity of the target speech act:
The final training set used for the released model contains 3,632 sentences/snippets, of which 985 (27%) are positive instances of disinformation/lie allegations. Class weighting (1:2.17) was applied to address the residual imbalance.
simpletransformers defaults.The final released model was trained on the entire 3,632-instance dataset (no evaluation split held out), after validation was completed via a separate 100-bootstrap cross-validation procedure (see below).
Performance was estimated via 100-bootstrap 80/20 train/evaluation splits:
| Version | MCC | Accuracy | F1 | AUROC | AUPRC |
|---|---|---|---|---|---|
| No class weights | 0.82 (0.03) | 0.93 (0.01) | 0.91 (0.01) | 0.97 (0.01) | 0.92 (0.02) |
| Weighted (1:2.17) | 0.82 (0.02) | 0.93 (0.01) | 0.91 (0.01) | 0.98 (0.01) | 0.92 (0.02) |
Values are means across 100 bootstrapped models; standard deviations in parentheses.
Usage example
from transformers import AutoModelForSequenceClassification, TextClassificationPipeline, AutoTokenizer, AutoConfig
MODEL = "mmochtak/lieline"
tokenizer = AutoTokenizer.from_pretrained(MODEL)
config = AutoConfig.from_pretrained(MODEL)
model = AutoModelForSequenceClassification.from_pretrained(MODEL)
pipe = TextClassificationPipeline(model=model, tokenizer=tokenizer, task='classification', device=0)
result = pipe([
"You are a moron.",
"You, sir, are a liar.",
"I do not know; I was not there.",
"You are giving us very misleading information!"
])
print(result)
Please cite the model as follows:
@misc{parlasent-model,
author = {Mochtak, Michal and Meijers, Maurits},
title = {LieLine},
year = {2026},
url = {https://huggingface.co/mmochtak/lieline},
publisher = {Hugging Face}
}