Downloads · 30 days
33
100% of all-time downloads
StanfordSCALE/assertion_sentence_has_answer_to_a_math_problem
assertion_sentence_has_answer_to_a_math_problem is a text classification model from StanfordSCALE. Use it when you need a label for a piece of text. It is set up for setfit.
This classifier was trained for EduBehaviors: Assertion-based schemas for auditable dialogue coding and is usable through the Python package EduBehaviors-kit. This classifier was trained on an LLM-annotated subset of…
Downloads · 30 days
33
100% of all-time downloads
All-time downloads
33
Public
Parameters
109M
438 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
This classifier was trained for EduBehaviors: Assertion-based schemas for auditable dialogue coding and is usable through the Python package EduBehaviors-kit. This classifier was trained on an LLM-annotated subset of teacher utterances from the TalkMoves Dataset. See the Datasets section below for more information.
| Dataset | Split | Size |
|---|---|---|
| StanfordSCALE/assertions_llm_annotated_talkmoves | train | 3,430 (53.3%) |
| StanfordSCALE/assertions_llm_annotated_talkmoves | dev | 858 (13.3%) |
| StanfordSCALE/assertions_llm_annotated_talkmoves | test | 2,146 (33.4%) |
This model's columns are assertion_sentence_has_answer_to_a_math_problem and split_sentence_has_answer_to_a_math_problem.
Base rate (share of rows labeled as True): 4.7% overall — 4.8% train, 5.1% dev, 4.4% test.
Labels were generated with LLM annotators. Krippendorff's alpha for this assertion is 0.409.
| Parameter | Value |
|---|---|
| Base model (body) | sentence-transformers/paraphrase-mpnet-base-v2 |
| Head | LogisticRegression |
| Body learning rate | 2e-05 |
| Head learning rate | 0.01 |
| Batch size | 16 (contrastive phase) / 32 (head) |
| Epochs | 10 |
| Max steps | 5000 (contrastive phase) |
| Eval max steps | 100 |
| Seed | 20260904 |
| Mixed precision | enabled on GPU |
| Split | n | Base rate | Precision | Recall | F1 (positive class) | ROC-AUC | Average precision |
|---|---|---|---|---|---|---|---|
| dev | 858 | 5.1% | 0.500 | 0.364 | 0.421 | 0.878 | 0.450 |
| test | 2,146 | 4.4% | 0.450 | 0.479 | 0.464 | 0.864 | 0.411 |
The model was trained on text built as:
{utterance}
The utterance is passed through as-is.
pip install setfit
from setfit import SetFitModel
model = SetFitModel.from_pretrained("StanfordSCALE/assertion_sentence_has_answer_to_a_math_problem")
text = 'Okay so thats the largest Yeah'
model.predict([text]) # -> array([1]) when the assertion holds
model.predict_proba([text]) # -> [[P(no), P(yes)]]
@misc{assertion_sentence_has_answer_to_a_math_problem,
author = {Stanford SCALE Initiative},
title = {Assertion classifier: sentence has answer to a math problem},
year = {2026},
url = {https://huggingface.co/StanfordSCALE/assertion_sentence_has_answer_to_a_math_problem}
}