Downloads · 30 days
0
Datadave09/DA-RoBERTa
DA-RoBERTa is a machine learning model from Datadave09. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This model corresponds to the paper "A Domain-adaptive Pre-training Approach for Language Bias Detection in News" (Krieger et al.,2022): https://github.com/Media-Bias-Group/A-Domain-adaptive-Pre-training-Approach-for-…
Downloads · 30 days
0
Access
Public
Updated Mar 21, 2022
Repo size
944 MB
Likes
2
Public
Click a slice to open those files.
.bin501 MB · 100%
From the Hugging Face model README
This model corresponds to the paper "A Domain-adaptive Pre-training Approach for Language Bias Detection in News" (Krieger et al.,2022): https://github.com/Media-Bias-Group/A-Domain-adaptive-Pre-training-Approach-for-Language-BiasDetection-in-News
The model can be used for sequence classification tasks of biased and non-biased language in news and media. It is initialized with roberta-base weights and fine-tuned on the Wiki Neutrality Corpus (Pryzant et al., 2020). More details on the training setup and experiments can be found in our paper.
You can use the model with the PyTorch framework
#imports
!pip install transformers
!pip install openpyxl
import torch
import torch.nn as nn
import numpy as np
from transformers import RobertaTokenizer,RobertaModel
#define model class including binary classification layer
class RobertaClass(torch.nn.Module):
def __init__(self):
super(RobertaClass, self).__init__()
self.roberta = RobertaModel.from_pretrained("roberta-base")
self.vocab_transform = torch.nn.Linear(768, 768)
self.dropout = torch.nn.Dropout(0.2)
self.classifier1 = torch.nn.Linear(768,2)
def forward(self, input_ids, attention_mask):
output_1 = self.roberta(input_ids=input_ids, attention_mask=attention_mask)
hidden_state = output_1[0]
pooler = hidden_state[:, 0]
pooler = self.vocab_transform(pooler)
pooler = self.dropout(pooler)
output = self.classifier1(pooler)
return output
#load model parameters
weight_dict = torch.load('DA-Roberta.bin')
#initialize model with fine-tuned parameters
model = RobertaClass()
model.load_state_dict(weight_dict)
#exemplary bias classification with instance extracted from BABE dataset (Spinde et al.,2021)
tokenizer = RobertaTokenizer.from_pretrained('roberta-base')
inputs = tokenizer("A cop shoots a Black man, and a police union flexes its muscle", return_tensors="pt")
outputs = model(**inputs)
if int(torch.argmax(outputs)) == 1:
print("Biased")
else:
print("Non-biased")
@InProceedings{Krieger2022,
author={Krieger, David and Spinde, Timo and Ruas, Terry and Kulshrestha, Juhi and Gipp, Bela},
booktitle={2022 ACM/IEEE Joint Conference on Digital Libraries (JCDL)},
title={A Domain-adaptive Pre-training Appraoch for Language Bias Detection in News},
year={2022},
address = "Cologne,Germany"
}