Downloads · 30 days
34
1% of all-time downloads
s-nlp/rubert-base-corruption-detector
rubert-base-corruption-detector is a text classification model from s-nlp. Use it when you need a label for a piece of text. It is set up for transformers.
This is a model for evaluation of naturalness of short Russian texts. It has been trained to distinguish human-written texts from their corrupted versions.
Downloads · 30 days
34
1% of all-time downloads
All-time downloads
3.8K
Public
Repo size
2.1 GB
Likes
1
Public
Click a slice to open those files.
.bin712 MB · 99%
From the Hugging Face model README
This is a model for evaluation of naturalness of short Russian texts. It has been trained to distinguish human-written texts from their corrupted versions.
Corruption sources: random replacement, deletion, addition, shuffling, and re-inflection of words and characters, random changes of capitalization, round-trip translation, filling random gaps with T5 and RoBERTA models. For each original text, we sampled three corrupted texts, so the model is uniformly biased towards the unnatural label.
Data sources: web-corpora from the Leipzig collection (rus_news_2020_100K, rus_newscrawl-public_2018_100K, rus-ru_web-public_2019_100K, rus_wikipedia_2021_100K), comments from OK and Pikabu.
On our private test dataset, the model has achieved 40% rank correlation with human judgements of naturalness, which is higher than GPT perplexity, another popular fluency metric.