Downloads · 30 days
11
13% of all-time downloads
CarolGuga/mbart-neutralization
mbart-neutralization is a machine learning model from CarolGuga. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
11
13% of all-time downloads
All-time downloads
86
Public
Parameters
611M
14.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.4 GB · 99%
From the Hugging Face model README
This model is a fine-tuned version of facebook/mbart-large-50 on an unknown dataset. It achieves the following results on the evaluation set:
This model is a fine-tuned version of facebook/mbart-large-50, a multilingual sequence-to-sequence Transformer model, adapted for the task of Spanish gender neutralization. The goal of the model is to transform gender-marked Spanish sentences into gender-neutral reformulations, preserving meaning while reducing grammatical gender marking. This task can be framed as a monolingual translation problem (Spanish → neutral Spanish). The model was trained using the Hugging Face Transformers library and follows a standard encoder–decoder architecture with transfer learning from the pretrained mBART model. The resulting system performs controlled rewriting rather than translation between languages, making it suitable for experiments in:
This model is intended for:
The model was trained on the Spanish Gender Neutralization dataset available on Hugging Face: 👉 hackathon-pln-es/neutral-es This dataset contains pairs of aligned sentences:
The dataset already includes a predefined split:
The dataset is relatively small and designed mainly for educational and experimental purposes, not for large-scale production systems. Before training, the data was:
Evaluation was performed using the BLEU score (sacrebleu), a standard metric in machine translation.
The model was trained using the Hugging Face Trainer API for sequence-to-sequence learning. Training steps:
The model therefore learns to perform monolingual rewriting via multilingual translation architecture.
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Bleu | Gen Len |
|---|---|---|---|---|---|
| No log | 1.0 | 440 | 0.0246 | 98.2861 | 18.5729 |
| 0.2226 | 2.0 | 880 | 0.0138 | 98.4772 | 18.5104 |