Downloads · 30 days
10
33% of all-time downloads
Megyy/lexiguard
lexiguard is a text classification model from Megyy. Use it when you need a label for a piece of text. The card lists the license as mit.
LexiGuard is a multilingual multitask model designed to detect and classify offensive language, with a focus on misogyny, misandry, and toxicity levels in English. The model also supports Slovak, making it suitable fo…
Downloads · 30 days
10
33% of all-time downloads
All-time downloads
30
Public
Parameters
278M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 98%
From the Hugging Face model README
LexiGuard is a multilingual multitask model designed to detect and classify offensive language, with a focus on misogyny, misandry, and toxicity levels in English. The model also supports Slovak, making it suitable for multilingual analysis of social media content.
It performs dual classification:
The model is based on xlm-roberta-base and was fine-tuned on a custom dataset primarily in English, with additional annotated samples in Slovak.
xlm-roberta-basefrom transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("Megyy/lexiguard")
model = AutoModelForSequenceClassification.from_pretrained("Megyy/lexiguard")
text = "Women are useless in politics."
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
# outputs.logits contains predictions for both tasks
Note: The model has two output heads:
- Head 1: Category (misogyny/misandry/neutral)
- Head 2: Toxicity (low/medium/high)
Task 1 – Category Classification
0: Neutral1: Misogyny2: MisandryTask 2 – Toxicity Prediction
0: Low1: Medium2: Highpytorch_model.bin / model.safetensors: model weightsconfig.json: model configurationtokenizer.json, vocab.txt, etc.: tokenizer filesREADME.md: model cardIf you use this model in your work, please cite:
@bachelorsthesis{majercakova2025lexiguard,
title={LexiGuard: Offensive Language Detection in English and Slovak Social Media},
author={Magdalena Majercakova},
year={2025},
note={Bachelor's thesis, TUKE},
}
Developed by Magdaléna Majerčáková as part of a Bachelor's Thesis
Supervised by Ing. Zuzana Sokolová, PhD
Faculty of Electrical Engineering and Informatics, TUKE (2025)