Downloads · 30 days
166
1% of all-time downloads
yhavinga/gpt2-large-dutch
gpt2-large-dutch is a text generation model from yhavinga. Use it when you need the model to write or continue text. It is set up for transformers.
A GPT2 large model (762M parameters) trained from scratch on Dutch, with perplexity 15.1 on cleaned Dutch mC4.
Downloads · 30 days
166
1% of all-time downloads
All-time downloads
16.5K
Public
Parameters
812M
72.8 GB on disk
Likes
9
Public
Click a slice to open those files.
.bin3.1 GB · 33%
How the weights are stored.
F32774M · 95%
From the Hugging Face model README
A GPT2 large model (762M parameters) trained from scratch on Dutch, with perplexity 15.1 on cleaned Dutch mC4.
You can use this GPT2-model directly with a pipeline for text generation.
MODEL_DIR='yhavinga/gpt2-large-dutch'
from transformers import pipeline, GPT2Tokenizer, GPT2LMHeadModel
tokenizer = GPT2Tokenizer.from_pretrained(MODEL_DIR)
model = GPT2LMHeadModel.from_pretrained(MODEL_DIR)
generator = pipeline('text-generation', model, tokenizer=tokenizer)
generated_text = generator('Het eiland West-', max_length=100, do_sample=True, top_k=40, top_p=0.95, repetition_penalty=2.0))
"Het eiland West-" - "Terschelling wordt sinds jaar en dag bewoond door de mens. De mensen die in het huidige Terherne wonen doen er alles aan om hun dorp te behouden voor deze diersoort, namelijk; een natuurreservaat dat vooral bestaat uit hoge duinen met lage begroeing waar planten van vroeger worden afgewisseld (zoals wilde hyacinten)en waarop grassen groeien waarvan sommige soorten zeldzame vormen hebben ontwikkeld: duinlelie of blauwe bosbes zijn bijvoorbeeld bekend vanwege onder andere kleurmole"
This model was trained on of the full configuration (33B tokens) of
cleaned Dutch mC4,
which is the original mC4, except
TL;DR: yhavinga/gpt2-medium-dutch is the best model.
a/b in the step-column have been trained to step a of a total of b steps.| model | params | train seq len | ppl | loss | batch size | epochs | steps | optim | lr | duration | config | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| yhavinga/gpt-neo-125M-dutch | gpt neo | 125M | 512 | 20.9 | 3.04 | 128 | 1 | 190000/558608 | adam | 2.4e-3 | 1d 12h | full |
| yhavinga/gpt2-medium-dutch | gpt2 | 345M | 512 | 15.1 | 2.71 | 128 | 1 | 320000/520502 | adam | 8e-4 | 7d 2h | full |
| yhavinga/gpt2-large-dutch | gpt2 | 762M | 512 | 15.1 | 2.72 | 32 | 1 | 1100000/2082009 | adafactor | 3.3e-5 | 8d 15h | large |
| yhavinga/gpt-neo-1.3B-dutch | gpt neo | 1.3B | 512 | 16.0 | 2.77 | 16 | 1 | 960000/3049896 | adafactor | 5e-4 | 7d 11h | full |
This project would not have been possible without compute generously provided by Google through the TPU Research Cloud. The HuggingFace 🤗 ecosystem was also instrumental in most, if not all, parts of the training. The following repositories where helpful in setting up the TPU-VM, and training the models:
Created by Yeb Havinga