Downloads · 30 days
8
1% of all-time downloads
fzliu/writing-style-classifier
writing-style-classifier is a text classification model from fzliu. Use it when you need a label for a piece of text. The card lists the license as mit.
Downloads · 30 days
8
1% of all-time downloads
All-time downloads
709
Public
Parameters
396M
1.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.6 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of ModernBERT-large for binary classification of human-written vs. AI-generated text. It outputs a single probability, P(AI), indicating the likelihood that a given input was generated by a language model. Vanguard AI-text detector was part of the 2nd-place system at PAN-CLEF 2026's Voight-Kampff Generative AI Detection leaderboard.
This model was trained on approximately 1.1 million texts from three datasets: DACTYL 2.0, LLMTrace, and MAGA-Bench.
The model was evaluated against two leading open-source AI text detectors, Fakespot and Desklib, across ten benchmark datasets. Three of these (dactyl-v2.0, llm-trace-eng, maga) are in-distribution with respect to this model's training data; the remaining seven (beemo, coconuts, detectrl, dolly-cosmopedia, originalityai, realdet, uchicago) are out-of-distribution and were not seen during training.
All F1 scores are macro-averaged and computed at a decision threshold of P(AI) = 0.5.
| Dataset | AUROC | Macro-F1 |
|---|---|---|
| dactyl-v2.0 † | 0.9903 | 0.9709 |
| llm-trace-eng † | 0.9909 | 0.9661 |
| maga † | 0.9992 | 0.9901 |
| beemo | 0.8560 | 0.6973 |
| coconuts | 0.9792 | 0.8690 |
| detectrl | 0.9370 | 0.8849 |
| dolly-cosmopedia | 0.9958 | 0.8884 |
| originalityai | 0.9322 | 0.7776 |
| realdet | 0.9540 | 0.9177 |
| uchicago | 0.9784 | 0.9105 |
† In-distribution (training data overlap)
Out-of-distribution averages: AUROC 0.9475, Macro-F1 0.8493
| Model | AUROC | Macro-F1 |
|---|---|---|
| Fakespot | 0.9315 | 0.8029 |
| Desklib | 0.9213 | 0.7837 |
| Vanguard | 0.9475 | 0.8493 |
This model should not be used as the sole basis for high-stakes decisions such as academic penalties or employment actions, given the false positive/negative rates documented below. Performance also degrades on text distributions not represented in training data; see evaluation results.
@article{thorat2026panclef,
title={Team DACTYL at PAN 2026: Bayesian Data Mixing and Empirical X-risk Minimization for AI-text Detection},
author={Thorat, Shantanu},
journal={Working Notes of CLEF},
year={2026}
}