Downloads · 30 days
217
1% of all-time downloads
THUMT/mGPT
mGPT is a text generation model from THUMT. Use it when you need the model to write or continue text. It is set up for transformers.
mGPT is pre-trained on the mC4 dataset using a causal language modeling objective. It was introduced in this paper and first released on this page.
Downloads · 30 days
217
1% of all-time downloads
All-time downloads
27K
Public
Repo size
2.3 GB
Likes
8
Public
Click a slice to open those files.
.bin1.1 GB · 100%
From the Hugging Face model README
mGPT is pre-trained on the mC4 dataset using a causal language modeling objective. It was introduced in this paper and first released on this page.
mGPT is a Transformer-based model which pre-trained on massive multilingual data covering over 101 languages. Similar to GPT-2, It was pre-trained on the raw texts only, with no human labeling. We use the same tokenization and vocabulary as the mT5 model.
You can use the raw model for text generation or using prompts for adapting it to a downstream task.
You can use this model directly with a pipeline for text generation. Here is how to use this model to get the features of a given text in PyTorch:
from transformers import MT5Tokenizer, GPT2LMHeadModel, TextGenerationPipeline
tokenizer = MT5Tokenizer.from_pretrained("THUMT/mGPT")
model = GPT2LMHeadModel.from_pretrained("THUMT/mGPT")
pipeline = TextGenerationPipeline(model=model, tokenizer=tokenizer)
text = "Replace me by any text you'd like."
text = pipeline(text, do_sample=True, max_length=1024)[0]["generated_text"]
The texts are tokenized using sentencepiece and a vocabulary size of 250,100. The inputs are sequences of 1,024 consecutive tokens. We use <extra_id_0> to separate lines in a document.
@misc{tan2021msp,
title={MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators},
author={Zhixing Tan and Xiangwen Zhang and Shuo Wang and Yang Liu},
year={2021},
eprint={2110.06609},
archivePrefix={arXiv},
primaryClass={cs.CL}
}