Downloads · 30 days
13
2% of all-time downloads
yanaiela/roberta-base-epoch_7
roberta-base-epoch_7 is a fill-mask model from yanaiela. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as mit.
This model is part of our reimplementation of the RoBERTa model, trained on Wikipedia and the Book Corpus only. We train this model for almost 100K steps, corresponding to 83 epochs. We provide the 84 checkpoints (inc…
Downloads · 30 days
13
2% of all-time downloads
All-time downloads
712
Public
Repo size
998 MB
Likes
0
Public
Click a slice to open those files.
.bin499 MB · 99%
From the Hugging Face model README
This model is part of our reimplementation of the RoBERTa model, trained on Wikipedia and the Book Corpus only. We train this model for almost 100K steps, corresponding to 83 epochs. We provide the 84 checkpoints (including the randomly initialized weights before the training) to provide the ability to study the training dynamics of such models, and other possible use-cases.
These models were trained in part of a work that studies how simple statistics from data, such as co-occurrences affects model predictions, which are described in the paper Measuring Causal Effects of Data Statistics on Language Model's `Factual' Predictions.
This is RoBERTa-base epoch_7.
This model was captured during a reproduction of RoBERTa-base, for English: it is a Transformers model pretrained on a large corpus of English data, using the Masked Language Modelling (MLM).
The intended uses, limitations, training data and training procedure for the fully trained model are similar to RoBERTa-base. Two major differences with the original model:
Using code from RoBERTa-base, here is an example based on PyTorch:
from transformers import pipeline
model = pipeline("fill-mask", model='yanaiela/roberta-base-epoch_83', device=-1, top_k=10)
model("Hello, I'm the <mask> RoBERTa-base language model")
@article{2207.14251,
Author = {Yanai Elazar and Nora Kassner and Shauli Ravfogel and Amir Feder and Abhilasha Ravichander and Marius Mosbach and Yonatan Belinkov and Hinrich Schütze and Yoav Goldberg},
Title = {Measuring Causal Effects of Data Statistics on Language Model's `Factual' Predictions},
Year = {2022},
Eprint = {arXiv:2207.14251},
}