Downloads · 30 days
36
92% of all-time downloads
ISAAC-corpus/isaac-generalization-segmentation
isaac-generalization-segmentation is a token classification model from ISAAC-corpus. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as cc-by-4.0.
Same weights as the public DiSCo release. This repository is the ISAAC-facing copy of BabakScrapes/disco-clause-segmenter; the checkpoints are identical. That card carries the full model description and the DiSCo corp…
Downloads · 30 days
36
92% of all-time downloads
All-time downloads
39
Public
Parameters
124M
496 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors496 MB · 99%
How the weights are stored.
F32124M · 100%
From the Hugging Face model README
Same weights as the public DiSCo release. This repository is the ISAAC-facing copy of
BabakScrapes/disco-clause-segmenter; the checkpoints are identical. That card carries the full model description and the DiSCo corpus details. This one documents how the model is used inside the ISAAC pipeline and restates the performance figures reported in the ISAAC manuscript. Cite whichever matches your use; if you are using ISAAC labels, cite both.
Token classifier that splits English text into clauses, the unit the situation-entity classifier then labels. Together the two models produce the clause-level generalization columns of the Illinois Social Attitudes Aggregate Corpus (ISAAC).
The model emits one tag per word (majority-voted across sub-word tokens). The
decoder in code/label_generalization.py reads them as follows:
2 marks a clause-final word; it closes the clause it appears in;0 and 1 mark clause-internal words;2;1.Reference implementation, including the sub-word majority vote:
import torch
from transformers import AutoModelForTokenClassification, AutoTokenizer
REPO = "ISAAC-corpus/isaac-generalization-segmentation"
tokenizer = AutoTokenizer.from_pretrained("roberta-base", use_fast=True,
add_prefix_space=True)
model = AutoModelForTokenClassification.from_pretrained(REPO).eval()
text = "My gay neighbor watered my plants while I was traveling"
words = text.split()
enc = tokenizer(words, is_split_into_words=True, return_tensors="pt",
truncation=True, max_length=512, padding="max_length")
word_ids = enc.word_ids(batch_index=0)
with torch.no_grad():
token_preds = model(**enc).logits[0].argmax(dim=-1).tolist()
per_word = [[] for _ in words]
for token_idx, word_id in enumerate(word_ids):
if word_id is not None:
per_word[word_id].append(token_preds[token_idx])
tags = [max(set(p), key=p.count) if p else 1 for p in per_word]
clauses, current, prev = [], [], 2
for word, tag in zip(words, tags):
if prev == 2:
current = []
current.append(word)
if tag == 2 and prev in (0, 1):
clauses.append(" ".join(current))
current = []
prev = tag
if current:
clauses.append(" ".join(current))
print(clauses)
Texts longer than 200 words are split on sentence boundaries before segmentation
in the ISAAC pipeline; see generalization.py in the Space.
The DiSCo corpus of opinionated, mixed-register English text (Hemmatian, 2022), with human-verified clause boundaries. See the DiSCo card.
Two FacebookAI/roberta-base models run in sequence: a clause segmenter, then a
18-way situation-entity classifier. Both are the same weights as
the public DiSCo release; see the
clause segmenter and
situation-entity classifier cards for the full model
description and the corpus they were trained on.
Segmentation. The segmenter covered most of the target clause span for 95.5% of human-verified clauses.
Classification, on the held-out 10% of the disco gold corpus (k ≈ 2,357):
| Target | Macro F1 | Accuracy |
|---|---|---|
| Full situation entity (18-way) | .514 | .737 |
| Genericity (2-way) | .852 | .860 |
| Eventivity (2-way) | .879 | .894 |
| Boundedness / habituality (4-way) | .804 | .850 |
The 18-way macro F1 of .514 is held down by heavy label imbalance across the rarer situation-entity types. The three collapsed features are what ISAAC actually reports, and they are the numbers to rely on.
Clause segmentation as the first stage of the ISAAC generalization pipeline.
Not a general-purpose syntactic parser, constituency parser, or sentence splitter. It targets the specific clause unit the situation-entity framework requires, which does not always coincide with a syntactic clause. English only.
| Try it without code | ISAAC Text Classifiers Space |
| Pipeline source, keyword lists, pattern sets | GitHub |
| Corpus download, samples, SQL playground | https://isaac.psychology.illinois.edu/ |
| Data Use Agreement | Data_Use_Agreement.md |
| Questions about the models or the corpus | isaac.corpus.support@gmail.com |
Please cite the ISAAC paper. One citation covers the whole project: the corpus, the pipeline, and every model. Please do not cite this model repository separately; keeping references in one place is what allows the project's citations to be found together.
@article{hemmatian2026isaac,
author = {Hemmatian, Babak and Hadjarab, Sarah and Chen, Jessica and Kurdi, Benedek},
title = {The {Illinois} Social Attitudes Aggregate Corpus ({ISAAC}): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale},
year = {2026},
journal = {arXiv},
eprint = {2609.27059},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
doi = {10.48550/arXiv.2609.27059},
url = {https://arxiv.org/abs/2609.27059}
}
Released under a Creative Commons Attribution 4.0 International License. You may use, share, and adapt these weights, including commercially, provided you give appropriate credit; see Citation above.
The ISAAC corpus itself is governed separately by the project Data Use Agreement.