Downloads · 30 days
8
5% of all-time downloads
rebego/mbart-neutralization
mbart-neutralization is a machine learning model from rebego. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of facebook/mbart-large-50 on the hackathon-pln-es/neutral-es dataset.
Downloads · 30 days
8
5% of all-time downloads
All-time downloads
149
Public
Parameters
611M
4.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.4 GB · 99%
From the Hugging Face model README
This model is a fine-tuned version of facebook/mbart-large-50 on the hackathon-pln-es/neutral-es dataset.
It learns to paraphrase gender-marked expressions into an inclusive style. For example, "La enfermera me curó" → "El personal sanitario me curó", thereby promoting more inclusive language.
It achieves the following results on the evaluation set:
mBART-50 is a pretrained multilingual encoder–decoder (sequence-to-sequence) model that supports 50 languages. It was designed to show that, instead of fine-tuning a separate model for each language pair, a single pre-trained model can be fine-tuned simultaneously on multiple translation directions. Building on the original mBART, it extends coverage by adding 25 more languages (for a total of 50), delivering a truly multilingual solution.
During pre-training, mBART-50 employs a denoising autoencoding objective: monolingual sentences are “noised” by randomly shuffling their order and span-masking a portion of tokens, and the model learns to reconstruct the original text.
Reducing gender bias in Spanish texts via monolingual style transfer.
Preprocessing step in NLP pipelines (e.g. for editorial tools or inclusive content generation).
As a basis for further fine-tuning on related sequence-to-sequence tasks (summarization, paraphrasing).
Only neutralizes gendered expressions in Spanish; it does not translate between languages.
Quality may degrade on domain-specific or very technical texts outside the training distribution.
May occasionally produce ungrammatical or awkward phrasing when forced to alter rare word combinations.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Bleu | Gen Len |
|---|---|---|---|---|---|
| No log | 1.0 | 440 | 0.0151 | 88.2841 | 34.8125 |
| 0.2281 | 2.0 | 880 | 0.0118 | 63.5448 | 36.7604 |