Downloads · 30 days
91
7% of all-time downloads
monsoon-nlp/ar-seq2seq-gender-decoder
ar-seq2seq-gender-decoder is a text generation model from monsoon-nlp. Use it when you need the model to write or continue text. It is set up for transformers.
This is a seq2seq model (decoder half) to "flip" gender in first-person Arabic sentences. The model can augment your existing Arabic data, or generate counterfactuals to test a model's decisions (would changing the ge…
Downloads · 30 days
91
7% of all-time downloads
All-time downloads
1.3K
Public
Parameters
191M
1.5 GB on disk
Likes
1
Public
Click a slice to open those files.
.bin765 MB · 50%
How the weights are stored.
F32191M · 100%
From the Hugging Face model README
This is a seq2seq model (decoder half) to "flip" gender in first-person Arabic sentences. The model can augment your existing Arabic data, or generate counterfactuals to test a model's decisions (would changing the gender of the subject or speaker change output?).
Intended Examples:
People's names, gender pronouns, gendered words (father, mother), and many other values are currently unchanged by this model. Future versions may be trained on more data.
import torch
from transformers import AutoTokenizer, EncoderDecoderModel
model = EncoderDecoderModel.from_encoder_decoder_pretrained(
"monsoon-nlp/ar-seq2seq-gender-encoder",
"monsoon-nlp/ar-seq2seq-gender-decoder",
min_length=40
)
tokenizer = AutoTokenizer.from_pretrained('monsoon-nlp/ar-seq2seq-gender-decoder') # same as MARBERT original
input_ids = torch.tensor(tokenizer.encode("أنا سعيدة")).unsqueeze(0)
generated = model.generate(input_ids, decoder_start_token_id=model.config.decoder.pad_token_id)
tokenizer.decode(generated.tolist()[0][1 : len(input_ids[0]) - 1])
> 'انا سعيد'
https://colab.research.google.com/drive/1S0kE_2WiV82JkqKik_sBW-0TUtzUVmrV?usp=sharing
I originally developed <a href="https://github.com/MonsoonNLP/el-la">a gender flip Python script</a> for Spanish sentences, using <a href="https://huggingface.co/dccuchile/bert-base-spanish-wwm-uncased">BETO</a>, and spaCy. More about this project: https://medium.com/ai-in-plain-english/gender-bias-in-spanish-bert-1f4d76780617
The Arabic model encoder and decoder started with weights and vocabulary from <a href="https://github.com/UBC-NLP/marbert">MARBERT from UBC-NLP</a>, and was trained on the <a href="https://camel.abudhabi.nyu.edu/arabic-parallel-gender-corpus/">Arabic Parallel Gender Corpus</a> from NYU Abu Dhabi. The text is first-person sentences from OpenSubtitles, with parallel gender-reinflected sentences generated by Arabic speakers.
Training notebook: https://colab.research.google.com/drive/1TuDfnV2gQ-WsDtHkF52jbn699bk6vJZV
This model is useful to generate male and female text samples, but falls short of capturing gender diversity in the world and in the Arabic language. This subject is discussed in the bias statement of the <a href="https://www.aclweb.org/anthology/2020.gebnlp-1.12/">Gender Reinflection paper</a>.