Downloads · 30 days
33
100% of all-time downloads
StanfordSCALE/assertion_sentence_has_comparison_terms
assertion_sentence_has_comparison_terms is a text classification model from StanfordSCALE. Use it when you need a label for a piece of text. It is set up for setfit.
This classifier was trained for EduBehaviors: Assertion-based schemas for auditable dialogue coding and is usable through the Python package EduBehaviors-kit. This classifier was trained on an LLM-annotated subset of…
Downloads · 30 days
33
100% of all-time downloads
All-time downloads
33
Public
Parameters
109M
438 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
This classifier was trained for EduBehaviors: Assertion-based schemas for auditable dialogue coding and is usable through the Python package EduBehaviors-kit. This classifier was trained on an LLM-annotated subset of teacher utterances from the TalkMoves Dataset. See the Datasets section below for more information.
| Dataset | Split | Size |
|---|---|---|
| StanfordSCALE/assertions_llm_annotated_talkmoves | train | 3,430 (53.3%) |
| StanfordSCALE/assertions_llm_annotated_talkmoves | dev | 858 (13.3%) |
| StanfordSCALE/assertions_llm_annotated_talkmoves | test | 2,146 (33.4%) |
This model's columns are assertion_sentence_has_comparison_terms and split_sentence_has_comparison_terms.
Base rate (share of rows labeled as True): 6.2% overall — 6.2% train, 6.3% dev, 6.2% test.
Labels were generated with LLM annotators. Krippendorff's alpha for this assertion is 0.657.
| Parameter | Value |
|---|---|
| Base model (body) | sentence-transformers/paraphrase-mpnet-base-v2 |
| Head | LogisticRegression |
| Body learning rate | 2e-05 |
| Head learning rate | 0.01 |
| Batch size | 16 (contrastive phase) / 32 (head) |
| Epochs | 10 |
| Max steps | 5000 (contrastive phase) |
| Eval max steps | 100 |
| Seed | 20260904 |
| Mixed precision | enabled on GPU |
| Split | n | Base rate | Precision | Recall | F1 (positive class) | ROC-AUC | Average precision |
|---|---|---|---|---|---|---|---|
| dev | 858 | 6.3% | 0.677 | 0.778 | 0.724 | 0.962 | 0.748 |
| test | 2,146 | 6.2% | 0.779 | 0.813 | 0.796 | 0.943 | 0.840 |
The model was trained on text built as:
{utterance}
The utterance is passed through as-is.
pip install setfit
from setfit import SetFitModel
model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_has_comparison_terms")
text = 'What I want you to focus on today is how can you relate the two diameters to the slant height and then how can you kind of think about the relationship between all three measurements'
model.predict([text]) # -> array([1]) when the assertion holds
model.predict_proba([text]) # -> [[P(no), P(yes)]]
@misc{assertion_sentence_has_comparison_terms,
author = {Stanford SCALE Initiative},
title = {Assertion classifier: sentence has comparison terms},
year = {2026},
url = {https://huggingface.co/StanfordSCALE/assertion_sentence_has_comparison_terms}
}