Downloads · 30 days
35
11% of all-time downloads
Tanor/SRGPTSENTPOS2
SRGPTSENTPOS2 is a text classification model from Tanor. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as cc-by-sa-4.0.
This is a positive-polarity classifier for Serbian WordNet synset glosses, fine-tuned from the GPT2-Orao model family. It is a component of the S6 sentiment-lexicon construction method described in:
Downloads · 30 days
35
11% of all-time downloads
All-time downloads
313
Public
Repo size
64.9 GB
Likes
0
Public
Click a slice to open those files.
.bin3.1 GB · 100%
From the Hugging Face model README
This is a positive-polarity classifier for Serbian WordNet synset glosses, fine-tuned from the GPT2-Orao model family. It is a component of the S6 sentiment-lexicon construction method described in:
Saša Petalinkar, Ranka M. Stanković, and Milica Ikonić Nešić (2025). Comparative analysis of methods for creating a sentiment lexicon of the Serbian WordNet. The Electronic Library, 43(4), 547–577. Paper and DOI.
Companion code and sentiment lexicons · Paired classifier · Original base model
| Field | Value |
|---|---|
| Model ID | Tanor/SRGPTSENTPOS2 |
| Architecture | GPT2ForSequenceClassification |
| Original base model | jerteh/gpt2-orao |
| Input | A Serbian synset gloss, representing one lexical meaning |
| Target | Positive vs non-positive polarity |
| Dataset expansion iteration | 2 (T2) |
| Label 0 | NON-POSITIVE |
| Label 1 | POSITIVE |
| Derived lexicon family | S6 |
| Paired classifier, same iteration | Tanor/SRGPTSENTNEG2 |
The linked training script initializes the classifier from jerteh/gpt2-orao. SRGPT is the project naming convention for the GPT2-Orao sentiment classifiers.
A non-positive label is the complement of the target class. It does not by itself mean that the gloss has the opposite polarity. A separate classifier handles that polarity.
The paper constructs polarity-labeled synsets from Serbian WordNet, starting from curated positive, negative, and objective seed sets and expanding through semantic relations. The initial sets reported in the paper contain 149 positive, 219 negative, and 19,475 objective synsets. Polarity-preserving relations expand the corresponding set; antonymy contributes to the opposite polarity. The selected datasets are T0, T2, T4, and T6, after zero, two, four, and six expansion iterations.
This checkpoint is associated with T2 and the POS classification task. Its numerical suffix is a dataset-expansion iteration, not an epoch count or a lexicon identifier.
The training script reads Sysnet from X_train_UPPOS2.csv and the target POS from y_train_UPPOS2.csv, replaces missing text with an empty string, and creates a stratified validation subset of 20% of that training CSV, with random_state=42. The UP inputs are the non-lemmatized gloss variant in the dataset-generation script.
The paper and code snapshot have different split descriptions: the paper describes an 80/10/10 split for neural models, while the scripts split existing training CSV files and create_sets.py leaves the initial split size at the library default. Exact checkpoint-specific sample assignments are not supplied in the public revision. The train_sets/ files referenced by the scripts are absent from that revision. Reconstructing the experiment requires the relevant lexical resources and saved preparation/split information; the paper's proportions alone do not establish this checkpoint's split.
For each iteration, the POS model estimates positive-class probability p_pos and the NEG model estimates negative-class probability p_neg. The lexicon-building code combines them as:
POS = p_pos * (1 - p_neg)
NEG = p_neg * (1 - p_pos)
OBJ = 1 - POS - NEG
It averages each score across the four iteration-specific pairs. A single checkpoint is one component of this construction; its two class probabilities are not the final three lexicon scores.
| Iteration | POS classifier | NEG classifier |
|---|---|---|
| 0 | SRGPTSENTPOS0 | SRGPTSENTNEG0 |
| 2 | SRGPTSENTPOS2 | SRGPTSENTNEG2 |
| 4 | SRGPTSENTPOS4 | SRGPTSENTNEG4 |
| 6 | SRGPTSENTPOS6 | SRGPTSENTNEG6 |
This example reads one gloss and returns both class probabilities. It pins the model to the weights revision present before the documentation update.
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
MODEL_ID = "Tanor/SRGPTSENTPOS2"
WEIGHTS_REVISION = "ac6d1549e49915ee2f1d771016f588c2ce61905d"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, revision=WEIGHTS_REVISION)
model = AutoModelForSequenceClassification.from_pretrained(
MODEL_ID, revision=WEIGHTS_REVISION
)
model.eval()
# An illustrative gloss, not a benchmark item or a claimed prediction.
inputs = tokenizer(
"koji oseća radost i zadovoljstvo",
return_tensors="pt", truncation=True, max_length=300,
)
with torch.inference_mode():
probabilities = model(**inputs).logits.softmax(dim=-1)[0]
scores = {model.config.id2label[i]: float(p) for i, p in enumerate(probabilities)}
print(scores)
The model uses standard PyTorch/Transformers sequence-classification classes, without custom remote code. The example requires PyTorch and Transformers. Where a Trainer record is available below, it includes software versions from the original run. No example prediction or benchmark score was generated for this documentation update.
The companion repository contains an evaluation report for SRGPT, POS, T2. It is preserved below, with class 0 = NON-POSITIVE and class 1 = POSITIVE. Metric values retain the source's rounding; support values are sample counts. The confusion matrix uses that class order.
[[4400 30]
[ 51 15]]
precision recall f1-score support
0 0.99 0.99 0.99 4430
1 0.33 0.23 0.27 66
accuracy 0.98 4496
macro avg 0.66 0.61 0.63 4496
weighted avg 0.98 0.98 0.98 4496
This is an archived experiment artifact, not a new evaluation. The report does not record a model-weight SHA, so its exact correspondence to the currently hosted weights has not been independently re-established. These binary-classification metrics are separate from evaluation of the derived lexicon and from sentiment evaluation of full sentences.
The following validation summary and training log are retained from the previous model card, without recomputing them. The recorded F1 is a training-validation metric, separate from the archived test report above and from lexicon or sentence-level sentiment evaluation. The source training code uses binary F1 for target label 1 when eval="f1" is selected.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | F1 |
|---|---|---|---|---|
| 0.0314 | 1.0 | 2697 | 0.1674 | 0.5111 |
| 0.0236 | 2.0 | 5394 | 0.1687 | 0.4308 |
| 0.0407 | 3.0 | 8091 | 0.1571 | 0.4 |
| 0.0086 | 4.0 | 10788 | 0.3466 | 0.3437 |
| 0.007 | 5.0 | 13485 | 0.1982 | 0.3750 |
| 0.0091 | 6.0 | 16182 | 0.2023 | 0.3492 |
| 0.0071 | 7.0 | 18879 | 0.2477 | 0.3678 |
| 0.0075 | 8.0 | 21576 | 0.1556 | 0.3750 |
| 0.0089 | 9.0 | 24273 | 0.2156 | 0.3729 |
The Trainer record and the linked source script are separate provenance sources. In particular, the recorded optimizer may differ from the script's adafactor setting; the historical record is retained without asserting that the linked script reproduces that exact run.
These settings describe the linked code revision, not a replacement for the per-run Trainer record.
| Setting | Source-code value |
|---|---|
| Maximum tokenized input length | 300 |
| Learning rate | 2e-5 |
| Training batch size per device | 1 |
| Evaluation batch size per device | 1 |
| Gradient accumulation | 4 steps |
| Optimizer | Adafactor |
| Weight decay | 0.01 |
| Validation split seed | 42 |
| Evaluation and saving | Each epoch |
| Early stopping patience | 3 evaluation calls |
| Epoch budget in experiment notebooks | Up to 32 |
The paper reports early-stopping patience of 3. The linked family script also uses 3.
This fine-tuned model is distributed under Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) (cc-by-sa-4.0), matching the license declared for its original base model, jerteh/gpt2-orao. Read the full LICENSE and NOTICE for attribution and the description of the fine-tuning changes.
The license permits sharing and adaptation, including commercial use, subject to its terms. Give appropriate credit, link to the license, and indicate changes. When sharing adapted material, apply the same license or a license permitted by its ShareAlike provisions.
The model license covers the model distribution and accompanying documentation within the rights the licensors can grant. External lexical resources, training datasets, and companion code retain their own terms.
An earlier description of Serbian WordNet reports CC BY-NC terms for the downloadable resource (Developing and Maintaining a WordNet: Procedures and Tools, 2014). The terms of the exact SrpWN release used for this fine-tuning, and any separate permission covering it, have not been verified in this documentation update. That historical statement alone does not determine the license of the trained weights. This model-license declaration does not grant rights to redistribute SrpWN or establish that every third-party permission for a particular use has been cleared. The remaining check is to identify the training release and record the applicable permission from its rights holders. See also Creative Commons guidance on AI training.
On 23 September 2026, the model-card declaration was changed from apache-2.0 to cc-by-sa-4.0 to align with the declared license of GPT2-Orao. The earlier declaration remains in the repository history; this update does not purport to revoke any valid rights previously granted.
When using this model family or the resulting lexicon-construction method, cite the paper and record the model ID and revision used.
@article{petalinkar2025sentimentlexicon,
author = {Petalinkar, Saša and Stanković, Ranka M. and Ikonić Nešić, Milica},
title = {Comparative analysis of methods for creating a sentiment lexicon of the Serbian WordNet},
journal = {The Electronic Library},
year = {2025},
volume = {43},
number = {4},
pages = {547--577},
doi = {10.1108/EL-08-2024-0253},
url = {https://doi.org/10.1108/EL-08-2024-0253}
}
ac6d1549e49915ee2f1d771016f588c2ce61905d. This pins the hosted checkpoint; it does not prove which weights produced every table in the paper.833582dbbf561a902fc5b872248db12bf529b3d9.