Downloads · 30 days
6
19% of all-time downloads
gokuls/bert_12_layer_model_v2_complete_training_new_wt_init
bert_12_layer_model_v2_complete_training_new_wt_init is a fill-mask model from gokuls. Use it when you need the model to fill a missing word. It is set up for transformers.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
6
19% of all-time downloads
All-time downloads
31
Public
Repo size
5.2 GB
Likes
0
Public
Click a slice to open those files.
.bin477 MB · 100%
From the Hugging Face model README
This model is a fine-tuned version of on the None dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Accuracy |
|---|---|---|---|---|
| 6.4002 | 0.08 | 10000 | 6.3571 | 0.1312 |
| 5.5302 | 0.16 | 20000 | 5.1885 | 0.2427 |
| 4.04 | 0.25 | 30000 | 3.8071 | 0.3863 |
| 3.7185 | 0.33 | 40000 | 3.4770 | 0.4246 |
| 3.5317 | 0.41 | 50000 | 3.3049 | 0.4441 |
| 3.4184 | 0.49 | 60000 | 3.1983 | 0.4558 |
| 3.3161 | 0.57 | 70000 | 3.1219 | 0.4650 |
| 3.2417 | 0.66 | 80000 | 3.0511 | 0.4726 |
| 3.1771 | 0.74 | 90000 | 2.9934 | 0.4789 |
| 3.1276 | 0.82 | 100000 | 2.9450 | 0.4850 |
| 3.0795 | 0.9 | 110000 | 2.9056 | 0.4895 |