Downloads · 30 days
20
8% of all-time downloads
minnesotanlp/scholawrite-bert-classifier
scholawrite-bert-classifier is a text classification model from minnesotanlp. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
This model is refered as BERT-SW-CLF in the paper. It is fined-tuned based on base-base-uncased Hugging Face, using train split of ScholaWrite dataset. The sole purpose of this model is to predict the next writing int…
Downloads · 30 days
20
8% of all-time downloads
All-time downloads
249
Public
Parameters
109M
438 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
This model is refered as BERT-SW-CLF in the paper. It is fined-tuned based on base-base-uncased Hugging Face, using train split of ScholaWrite dataset. The sole purpose of this model is to predict the next writing intention given scholarly writing in latex.
The model is intended to used for next writing intention prediction in LaTex paper draft. It takes 'before' text warped by special tokens as input, and output the next writing intention which is 1 of 15 predefined labels.
The model is fine-tuned only for next writing intention prediction and infereneced in closed enviroment. Its main goal is to examine the usefullness of our dataset. It is suitable for acdamic use, but not suitable for production, general public use, or consumer-oriented service. In addition, use this model on tasks besides next intention prediction in LaTex paper draft may not work well.
The bias and limitations of this model mainly came from the dataset (<span style="font-variant: small-caps;">ScholaWrite</span>) it fine-tuned on.
First, the <span style="font-variant: small-caps;">ScholaWrite</span> dataset is currently limited to the computer science domain, as LaTeX is predominantly used in computer science journals and conferences. This domain-specific focus in dataset may restrict the model's generalizability to other scientific disciplines. Future work could address this limitation by collecting keystroke data from a broader range of fields with diverse writing conven554 tions and tools, such as the humanities or biological sciences. For example, students in humanities usu556 ally write book-length papers and integrate more sources, so it could affect cognitive complexities.
Second, all participants were early-career researchers (e.g., PhD students) at an R1 university in the United States, which means the models may not learn the professional writing behavior and cognitive process from expert. Expanding the dataset to include senior researchers, such as post-doctoral fellows and professors, could offer valuable insights into how writing strategies and revision behaviors evolve with research experience and expertise.
Third, the dataset is exclusive to English-language writing, which restricts model's capability to predict next writing intention in multilingual or non-English contexts. Expanding to multilingual settings could reveal unique cognitive and linguistic insights into writing across languages.
import os
from dotenv import load_dotenv
import torch
from transformers import BertTokenizer, BertForSequenceClassification, RobertaTokenizer, RobertaForSequenceClassification
from huggingface_hub import login
load_dotenv()
HUGGINGFACE_TOKEN = os.getenv("HUGGINGFACE_TOKEN")
login(token=HUGGINGFACE_TOKEN)
TOTAL_CLASSES = 15
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
tokenizer.add_tokens("<INPUT>") # start input
tokenizer.add_tokens("</INPUT>") # end input
tokenizer.add_tokens("<BT>") # before text
tokenizer.add_tokens("</BT>") # before text
tokenizer.add_tokens("<PWA>") # start previous writing action
tokenizer.add_tokens("</PWA>") # end previous writing action
model = BertForSequenceClassification.from_pretrained('minnesotanlp/scholawrite-bert-classifier', num_labels=TOTAL_CLASSES)
before_text = "sample before text"
text = "<INPUT>" + "<BT>" + before_text + "</BF> " + "</INPUT>"
input = tokenizer(text, return_tensors="pt")
pred = model(input["input_ids"]).logits.argmax(1)
print("class:", pred)
This model is fine-tuned on minnesotanlp/scholawrite dataset train split. It is keystroke logs of an end-to-end scholarly writing process, with thorough annotations of cognitive writing intentions behind each keystroke. No additional data pre-processing or filtering performed on the dataset.
The model was fine tuned by passing in the before_text section of a prompt as the input, and using the intention as the ground truth data. The model output an integer according to each intention label (1-15).
The data has class imbalanced on both training and testing data splits, so we use weighted F1 to measure the performance.
| BERT | RoBERTa | LLama-8B-Instruct | GPT-4o | |
|---|---|---|---|---|
| Base | 0.04 | 0.02 | 0.12 | 0.08 |
| + SW | 0.64 | 0.64 | 0.13 | - |
Table above presents the weighted F1 scores for predicting writing intentions across baselines and fine-tuned models. All models finetuned on ScholaWrite show a improvement performance compared to their baselines. BERT and RoBERTa achieved the most improvement, while LLama-8B-Instruct showed a modest improvement after fine-tuning. Those results demonstrate the effectiveness of our ScholaWrite dataset to align language models with writers' intentions.
@misc{wang2025scholawritedatasetendtoendscholarly,
title={ScholaWrite: A Dataset of End-to-End Scholarly Writing Process},
author={Linghe Wang and Minhwa Lee and Ross Volkov and Luan Tuyen Chau and Dongyeop Kang},
year={2025},
eprint={2502.02904},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2502.02904},
}