Downloads · 30 days
0
gameresearch/modernbert-eie-2048
modernbert-eie-2048 is a text classification model from gameresearch. Use it when you need a label for a piece of text. The card lists the license as apache-2.0.
This model is a fine tuned ModernBERT base encoder for binary text classification of Steam reviews, predicting whether a review expresses an Emotionally Impactful Experience (EIE). It outputs a probability (and label)…
Downloads · 30 days
0
Access
Public
Updated Jul 22, 2026
Repo size
598 MB
Likes
0
Public
Click a slice to open those files.
.safetensors598 MB · 99%
From the Hugging Face model README
This model is a fine tuned ModernBERT base encoder for binary text classification of Steam reviews, predicting whether a review expresses an Emotionally Impactful Experience (EIE). It outputs a probability (and label) for:
The model was designed to be scalable to millions of reviews and robust to long, messy, player written text.
This model is a fine-tuned ModernBERT-base supporting 2048-token inputs. Training used weighted cross-entropy to handle class imbalance (≈26% positives), AdamW with a conservative learning rate, and early stopping based on validation average precision (AUPRC). The best checkpoint by validation AUPRC was reloaded and saved for release. The model targets robust performance across long-text reviews.
Binary classification of review-like texts. Positive class corresponds to an expression of an emotionally impactful experience, which includes emotionally moving/challengin/discomforting experiences. See associated paper for a more detailed discussion and definition of the concept and review annotation. The model outputsa logit/probability of EIE and a binary label using an optimized threshold (selected to maximize F1 on validation data).
Always evaluate the model on your target domain.
Using HuggingFace Transformers (Python):
import torch
from transformers import AutoTokenizer, AutoConfig, AutoModelForSequenceClassification
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
base_model_id = "answerdotai/ModernBERT-base"
fine_tuned_model_id = "gameresearch/modernbert-eie-2048"
threshold = 0.2 # optimized decision threshold
tokenizer = AutoTokenizer.from_pretrained(
fine_tuned_model_id,
subfolder="model",
)
config = AutoConfig.from_pretrained(base_model_id)
config.num_labels = 2
model = AutoModelForSequenceClassification.from_config(config)
weights_filename = "model/model.safetensors"
weights_path = hf_hub_download(fine_tuned_model_id, weights_filename)
state_dict = load_file(weights_path)
model.load_state_dict(state_dict)
model.eval()
def predict_eie(review_text: str):
inputs = tokenizer(
review_text,
truncation=True,
padding="max_length",
max_length=2048,
return_tensors="pt",
)
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits.squeeze()
probs = torch.softmax(logits, dim=-1)
prob_eie = probs[1].item()
label = int(prob_eie >= threshold)
return {"prob_eie": prob_eie, "label": label}
print(predict_eie("This game absolutely destroyed me emotionally. I still think about the ending."))
</small>
The model was trained on 2,500 manually annotated Steam reviews (English). For greater details on the labeling scheme, sampling strategy and the annotation process check the associated paper.
No text cleaning was performed.
Objective:
A hyperparameter search over learning rates {1×10^-5, 2×10^-5, 3×10^-5} and epoch counts {3, 5} was conducted. Selection criteria was best validation AUPRC.
Optimizer and schedule:
Reproducibility:
Batching and precision:
Early stopping and checkpointing:
Threshold Tuning:
Labeled data split:
The test set remained locked until final evaluation.
Primary evaluation metrics:
All metrics are computed on the held out test set using the optimized decision threshold .
Test accuracy: 0.957 Test F1 score: 0.919 Test AUPRC: 0.964
Error analysis:
False negatives:
False positives:
The model shows high performance on the EIE classification task (F1 = 0.919, AUPRC = 0.964).
Misclassifications mainly occur on ambiguous, borderline cases, consistent with conceptual fuzziness of EIE and human annotation disagreements.
Error analysis found no clear systematic bias toward false positives or false negatives or specific keywords, beyond occasional over‑reliance on lexical cues.
Carbon emissions were measured with CodeCarbon (Courty et al., 2024) using measure_power_secs=1.
GPU: NVIDIA Quadro RTX 5000, 12 GB VRAM.
Anonymized