Downloads · 30 days
11
41% of all-time downloads
morten-j/pre-train_mBERT
pre-train_mBERT is a fill-mask model from morten-j. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
11
41% of all-time downloads
All-time downloads
27
Public
Parameters
178M
712 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors712 MB · 99%
From the Hugging Face model README
This model is a fine-tuned version of google-bert/bert-base-multilingual-cased on an unknown dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 1.4994 | 1.0 | 368814 | 1.3694 |
| 1.3718 | 2.0 | 737628 | 1.2540 |
| 1.2979 | 3.0 | 1106442 | 1.1986 |