Downloads · 30 days
9.4K
2% of all-time downloads
GottBERT/GottBERT_base_last
GottBERT_base_last is a fill-mask model from GottBERT. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as mit.
GottBERT is the first German-only RoBERTa model, pre-trained on the German portion of the first released OSCAR dataset. This model aims to provide enhanced natural language processing (NLP) performance for the German…
Downloads · 30 days
9.4K
2% of all-time downloads
All-time downloads
415K
Public
Parameters
127M
1.5 GB on disk
Likes
17
Public
Click a slice to open those files.
.bin507 MB · 33%
From the Hugging Face model README
GottBERT is the first German-only RoBERTa model, pre-trained on the German portion of the first released OSCAR dataset. This model aims to provide enhanced natural language processing (NLP) performance for the German language across various tasks, including Named Entity Recognition (NER), text classification, and natural language inference (NLI). GottBERT has been developed in two versions: a base model and a large model, tailored specifically for German-language tasks.
This was presented in GottBERT: a pure German Language Model.
GottBERT was evaluated across various downstream tasks:
Mertics:
Details:
| Model | Accuracy NLI | GermEval_14 F1 | CoNLL F1 | Coarse F1 | Fine F1 | 10kGNAD F1 |
|---|---|---|---|---|---|---|
| GottBERT_base_best | 80.82 | 87.55 | <ins>85.93</ins> | 78.17 | 53.30 | 89.64 |
| GottBERT_base_last | 81.04 | 87.48 | 85.61 | <ins>78.18</ins> | 53.92 | 90.27 |
| GottBERT_filtered_base_best | 80.56 | <ins>87.57</ins> | 86.14 | 78.65 | 52.82 | 89.79 |
| GottBERT_filtered_base_last | 80.74 | 87.59 | 85.66 | 78.08 | 52.39 | 89.92 |
| GELECTRA_base | 81.70 | 86.91 | 85.37 | 77.26 | 50.07 | 89.02 |
| GBERT_base | 80.06 | 87.24 | 85.16 | 77.37 | 51.51 | 90.30 |
| dbmdzBERT | 68.12 | 86.82 | 85.15 | 77.46 | 52.07 | 90.34 |
| GermanBERT | 78.16 | 86.53 | 83.87 | 74.81 | 47.78 | 90.18 |
| XLM-R_base | 79.76 | 86.14 | 84.46 | 77.13 | 50.54 | 89.81 |
| mBERT | 77.03 | 86.67 | 83.18 | 73.54 | 48.32 | 88.90 |
| GottBERT_large | 82.46 | 88.20 | <ins>86.78</ins> | 79.40 | 54.61 | 90.24 |
| GottBERT_filtered_large_best | 83.31 | 88.13 | 86.30 | 79.32 | 54.70 | 90.31 |
| GottBERT_filtered_large_last | 82.79 | <ins>88.27</ins> | 86.28 | 78.96 | 54.72 | 90.17 |
| GELECTRA_large | 86.33 | <ins>88.72</ins> | <ins>86.78</ins> | 81.28 | <ins>56.17</ins> | 90.97 |
| GBERT_large | <ins>84.21</ins> | <ins>88.72</ins> | 87.19 | <ins>80.84</ins> | 57.37 | <ins>90.74</ins> |
| XLM-R_large | 84.07 | 88.83 | 86.54 | 79.05 | 55.06 | 90.17 |
Get the fairseq checkpoints here.
If you use GottBERT in your research, please cite the following paper:
@inproceedings{scheible-etal-2024-gottbert,
title = "{G}ott{BERT}: a pure {G}erman Language Model",
author = "Scheible, Raphael and
Frei, Johann and
Thomczyk, Fabian and
He, Henry and
Tippmann, Patric and
Knaus, Jochen and
Jaravine, Victor and
Kramer, Frank and
Boeker, Martin",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year = "2024",
address = "Miami, Florida, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.emnlp-main.1183",
pages = "21237--21250",
}