Downloads · 30 days
15
31% of all-time downloads
asseco-group/roberta-incoherence-classifier
roberta-incoherence-classifier is a text classification model from asseco-group. Use it when you need a label for a piece of text. The card lists the license as cc-by-sa-4.0.
<h1 align="center"roberta-incoherence-classifier</h1
Downloads · 30 days
15
31% of all-time downloads
All-time downloads
48
Public
Parameters
443M
1.8 GB on disk
Likes
2
Trending 1
Click a slice to open those files.
.safetensors1.8 GB · 99%
From the Hugging Face model README
Encoder-based classifier for document inconsistency detection in Polish. This model evaluates the semantic consistency between two text fragments (e.g. sections of legal, procurement or organizational documents). It follows an NLI-like setup but redefines labels specifically for document coherence auditing. This model was initalized from PKOBP/polish-roberta-8k and adapted into an inconsistency classifier through supervised training on high-quality document-style pairs.
Not intended for:
Finetuning on specific domain data is recommended for best production accuracy.
| Label | Meaning |
|---|---|
| entailment | Hypothesis is a faithful, condensed or paraphrased restatement of the premise. All critical constraints, actors, conditions and scope remain intact. |
| neutral | Hypothesis neither follows nor contradicts the premise. Typically introduces unverifiable or out‑of‑scope information (e.g. different institutions, expanded context, unrelated assumptions). |
| contradiction | Hypothesis directly conflicts with the premise: reverses permissions/requirements, changes legal scope, numeric limits, formats, dates, or the responsible authority or both statements cannot realistically be true at the same time. |
Rule: A single critical mismatch (date / territory / authority / format / obligation vs. optional) is sufficient for contradiction, even if most of the text agrees.
asseco-group/roberta-incoherence-classifiergradient_accumulation_steps=112e-5, warmup ratio: 0.1, weight decay: 0.010.05 precision recall f1-score support
entailment 0.94 0.90 0.92 150
neutral 0.87 0.91 0.89 150
contradiction 0.93 0.93 0.93 150
accuracy 0.91 450
macro avg 0.91 0.91 0.91 450
weighted avg 0.91 0.91 0.91 450
While the task is NLI-like, the label semantics are redefined for document-level procedural consistency, for which no direct open-source baselines currently exist.
import torch
from transformers import pipeline
device = "cuda" if torch.cuda.is_available() else "cpu"
classifier = pipeline(
"text-classification",
model="asseco-group/roberta-incoherence-classifier",
tokenizer="asseco-group/roberta-incoherence-classifier",
top_k=None,
return_all_scores=True,
device=device
)
premise = (
"Wykonawca dostarczy pliki w formacie .shp zgodne z oprogramowaniem ArcGIS 10.2, "
"wraz z mapami wydrukowanymi w formacie A4."
)
hypo = (
"Wykonawca przekaże wyłącznie pliki .kml kompatybilne z QGIS "
"i przygotuje dokumentację w formacie A3."
)
result = classifier({"text": premise, "text_pair": hypo})
print(result)
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
name = "asseco-group/roberta-incoherence-classifier"
tokenizer = AutoTokenizer.from_pretrained(name, use_fast=True)
model = AutoModelForSequenceClassification.from_pretrained(name).eval()
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
pairs = [
("Zwrot kosztów w 60 dni ...", "Zwrot kosztów nastąpi w 30 dni ..."),
]
enc = tokenzier(
[p for p, h in pairs],
[h for p, h in pairs],
padding=True, truncation=True, max_length=512,
return_tensors="pt"
).to(device)
with torch.no_grad():
logits = model(**enc).logits
probs = logits.softmax(-1).cpu()
print(probs)
@misc{asseco2025incoherence,
title = {Polish RoBERTa-based Incoherence/Consistency Classifier (encoder-only)},
author = {Asseco Group},
year = {2025},
url = {https://huggingface.co/asseco-group/roberta-incoherence-classifier}
}