Downloads · 30 days
10
8% of all-time downloads
gplsi/Aitana-ClearLangDetection-R-1.0
Aitana-ClearLangDetection-R-1.0 is a text classification model from gplsi. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
- Model Description - Training Details - Technical Specifications - Evaluation - Additional Information
Downloads · 30 days
10
8% of all-time downloads
All-time downloads
129
Public
Parameters
283M
3.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt2.3 GB · 66%
From the Hugging Face model README
This model is fine-tuned from BSC-LT/mRoBERTa for the task of clear language classification in Spanish texts.
It predicts among three categories of linguistic clarity:
The dataset consists of Spanish texts annotated with clarity levels:
The original dataset is available in gplsi/discriminative_clearsim_es.
Confusion Matrix
| Pred FAC | Pred LF | Pred TXT | |
|---|---|---|---|
| True FAC | 1373 | 15 | 8 |
| True LF | 29 | 1367 | 0 |
| True TXT | 16 | 1 | 1379 |
| Class | Precision | Recall | F1-score | Support |
|---|---|---|---|---|
| FAC | 0.9683 | 0.9835 | 0.9758 | 1396 |
| LF | 0.9884 | 0.9792 | 0.9838 | 1396 |
| TXT | 0.9942 | 0.9878 | 0.9910 | 1396 |
Confusion Matrix
| Pred FAC | Pred LF | Pred TXT | |
|---|---|---|---|
| True FAC | 1220 | 13 | 8 |
| True LF | 28 | 1213 | 0 |
| True TXT | 13 | 1 | 1227 |
| Class | Precision | Recall | F1-score | Support |
|---|---|---|---|---|
| FAC | 0.9675 | 0.9831 | 0.9752 | 1241 |
| LF | 0.9886 | 0.9774 | 0.9830 | 1241 |
| TXT | 0.9935 | 0.9887 | 0.9911 | 1241 |
Confusion Matrix
| Pred FAC | Pred LF | Pred TXT | |
|---|---|---|---|
| True FAC | 153 | 2 | 0 |
| True LF | 1 | 154 | 0 |
| True TXT | 3 | 0 | 152 |
| Class | Precision | Recall | F1-score | Support |
|---|---|---|---|---|
| FAC | 0.9745 | 0.9871 | 0.9808 | 155 |
| LF | 0.9872 | 0.9936 | 0.9903 | 155 |
| TXT | 1.0000 | 0.9806 | 0.9902 | 155 |
For training, we used custom code developed to fine-tuned model using Transformers library.
This model was trained on NVIDIA DGX systems equipped with A100 GPUs, which enabled efficient large-scale training. For this model, we used one A100 GPU.
The model has been developed by the Language and Information Systems Group (GPLSI) and the Centro de Inteligencia Digital (CENID), both part of the University of Alicante (UA), as part of their ongoing research in Natural Language Processing (NLP).
This work is funded by the Ministerio para la Transformación Digital y de la Función Pública, co-financed by the EU – NextGenerationEU, within the framework of the project Desarrollo de Modelos ALIA.
We would like to express our gratitude to all individuals and institutions that have contributed to the development of this work.
Special thanks to:
We also acknowledge the financial, technical, and scientific support of the Ministerio para la Transformación Digital y de la Función Pública - Funded by EU – NextGenerationEU within the framework of the project Desarrollo de Modelos ALIA, whose contribution has been essential to the completion of this research.
This model has been developed and fine-tuned specifically for classification task. The authors are not responsible for potential errors, misinterpretations, or inappropriate use of the model beyond its intended purpose.
If you use this model in your research or work, please cite it as follows:
@misc{gplsi-aitama-clear-r-1.0,
author = {Sepúlveda-Torres, Robiert and Martínez-Murillo, Iván and Bonora, Mar and Consuegra-Ayala, Juan Pablo and Galeano, Santiago and Miró Maestre, María and and Grande, Eduardo and Canal-Esteve, Miquel and Estevanell-Valladares, Ernesto L. and Yáñez-Romero, Fabio and Gutierrez, Yoan and Abreu Salas, José Ignacio and Lloret, Elena and Montoyo, Andrés and Muñoz-Guillena and Palomar, Manuel},
title = {Aitana-ClearLangDetection-R-1.0: Fine-tuned model for clear language classification (TXT, FAC, LF)},
year = {2025},
institution = {Language and Information Systems Group (GPLSI) and Centro de Inteligencia Digital (CENID), University of Alicante (UA)},
howpublished = {\url{https://huggingface.co/gplsi/Aitana-ClearLangDetection-R-1.0}},
note = {Accessed: 2025-10-03}
}
Copyright © 2026 Language and Information Systems Group (GPLSI) and Centro de Inteligencia Digital (CENID), University of Alicante (UA). Distributed under the Apache License 2.0.