Downloads · 30 days
0
ltg/ssa-perin
ssa-perin is a token classification model from ltg. Use it when you need labels on individual words, such as names. The card lists the license as apache-2.0.
We here release a pretrained model (and an easy-to-run wrapper) for structured sentiment analysis (SSA) of Norwegian text, trained on the NoReCfine dataset. It implements a method described in the paper Direct parsing…
Downloads · 30 days
0
Access
Public
Updated Mar 8, 2024
Repo size
1.1 GB
Likes
1
Public
Click a slice to open those files.
.bin1.1 GB · 97%
From the Hugging Face model README
We here release a pretrained model (and an easy-to-run wrapper) for structured sentiment analysis (SSA) of Norwegian text, trained on the NoReC_fine dataset. It implements a method described in the paper Direct parsing to sentiment graphs by Samuel et al. 2022 which demonstrated how a graph-based semantic parser (PERIN) can be applied to the task of structured sentiment analysis, directly predicting sentiment graphs from text.
The model will attempt to identify the following components for a given sentence it deems to be sentiment-bearing: source expressions (the opinion holder), target expressions (what the opinion is directed towards), polar expressions (the part of the text indicating that an opinion is expressed), and finally the polarity (positive or negative). For more information about how these categories are defined in the training data, please see the paper A Fine-grained Sentiment Dataset for Norwegian by Øvrelid et al. 2020. For each identified expression, the character offsets in the text are also provided.
Here is an example showing how to use the model for predicting such sentiment tuples:
>>> import model_wrapper
>>> model = model_wrapper.PredictionModel()
>>> model.predict(['vi liker svart kaffe'])
[{'sent_id': '0',
'text': 'vi liker svart kaffe',
'opinions': [{'Source': [['vi'], ['0:2']],
'Target': [['svart', 'kaffe'], ['9:14', '15:20']],
'Polar_expression': [['liker'], ['3:8']],
'Polarity': 'Positive'}]}]
The model is trained on NoReC_fine, a dataset for fine-grained sentiment analysis in Norwegian, based on a subset of documents from the Norwegian Review Corpus (NoReC) which constists of professionally authored reviews from multiple news-sources and across a wide variety of domains, including literature, games, music, products, movies and more.
The method proposed by Samuel et al. (2022) suggests three different ways to encode sentiment graphs: "node-centric", "labeled-edge", and "opinion-tuple". The model released here uses the following configuration:
The model achieves the following results on the held-out test set of NoReC_fine (see the paper for description the metrics):
If you use this model in your academic work, please quote the following paper:
@inproceedings{samuel2022,
title={Direct parsing to sentiment graphs},
author={David Samuel and Jeremy Barnes and Robin Kurtz and
Stephan Oepen and Lilja Øvrelid and Erik Velldal},
year={2022},
booktitle = "Proceedings of the 60th Annual Meeting of
the Association for Computational Linguistics",
address = "Dublin, Ireland"
}
Erik Velldal and Larisa Kolesnichenko