Downloads · 30 days
17
2% of all-time downloads
vicclab/FolkGPT
FolkGPT is a text generation model from vicclab. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of gpt2 on vicclab/fairytales dataset.
Downloads · 30 days
17
2% of all-time downloads
All-time downloads
839
Public
Repo size
2.6 GB
Likes
2
Public
Click a slice to open those files.
.bin510 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of gpt2 on vicclab/fairy_tales dataset.
This model is the result of fine-tuning gpt2 on a dataset of fairy tales from various cultures.
The idea behind this is to generate text in the fashion of fairy tales written in the 18th and 19th centuries.
Why? Fairy tales seemed an appropriate application for text generation, as stories are usually short(ish), self-contained, and easy to read.
Trained on the vicclab/fairy_tales dataset. The dataset consists of a number of texts which were downloaded from Project Gutenberg, and then edited to remove all text except for the stories themselves. These were then all concatenated into a text file and pushed to HF at https://huggingface.co/datasets/vicclab/fairy_tales. The latest update to the dataset, which was used in the training of this model, was created and uploaded on February 26th, 2023. Texts used [and token count after removing boilerplate text]: https://www.gutenberg.org/files/2591/2591-0.txt [102927 tokens] https://www.gutenberg.org/files/503/503-0.txt [138353 tokens] https://www.gutenberg.org/cache/epub/69739/pg69739.txt [51035 tokens] https://www.gutenberg.org/files/2435/2435-0.txt [98791 tokens] https://www.gutenberg.org/cache/epub/7871/pg7871.txt [49410 tokens] https://www.gutenberg.org/files/8933/8933-0.txt [178622 tokens] gutenberg.org/cache/epub/30834/pg30834.txt [58359 tokens] https://www.gutenberg.org/cache/epub/68589/pg68589.txt [39815 tokens] https://www.gutenberg.org/cache/epub/34453/pg34453.txt [69365 tokens] gutenberg.org/cache/epub/8653/pg8653.txt [35351]
[Total tokens in actual dataset: 1002654 tokens]
The dataset was loaded, sampling by paragraph. From here, the dataset was split into a training dataset and a validation dataset in an 80-20 split. These were then tokenized. The model was set up, and the trainer was instantiated with the training_arguments listed below. Then, the training took place.
The following hyperparameters were used during training: