Downloads · 30 days
13
3% of all-time downloads
AndyChiang/cdgp-csg-roberta-dgen
cdgp-csg-roberta-dgen is a fill-mask model from AndyChiang. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as mit.
This model is a Candidate Set Generator in "CDGP: Automatic Cloze Distractor Generation based on Pre-trained Language Model", Findings of EMNLP 2022.
Downloads · 30 days
13
3% of all-time downloads
All-time downloads
454
Public
Repo size
998 MB
Likes
0
Public
Click a slice to open those files.
.bin499 MB · 99%
From the Hugging Face model README
This model is a Candidate Set Generator in "CDGP: Automatic Cloze Distractor Generation based on Pre-trained Language Model", Findings of EMNLP 2022.
Its input are stem and answer, and output is candidate set of distractors. It is fine-tuned by DGen dataset based on roberta-base model.
For more details, you can see our paper or GitHub.
from transformers import RobertaTokenizer, RobertaForMaskedLM, pipeline
tokenizer = RobertaTokenizer.from_pretrained("AndyChiang/cdgp-csg-roberta-dgen")
csg_model = RobertaForMaskedLM.from_pretrained("AndyChiang/cdgp-csg-roberta-dgen")
unmasker = pipeline("fill-mask", tokenizer=tokenizer, model=csg_model, top_k=10)
sent = "The only known planet with large amounts of water is <mask>. </s> earth"
cs = unmasker(sent)
print(cs)
This model is fine-tuned by DGen dataset, which covers multiple domains including science, vocabulary, common sense and trivia. It is compiled from a wide variety of datasets including SciQ, MCQL, AI2 Science Questions, etc. The detail of DGen dataset is shown below.
| DGen dataset | Train | Valid | Test | Total |
|---|---|---|---|---|
| Number of questions | 2321 | 300 | 259 | 2880 |
You can also use the dataset we have already cleaned.
We use a special way to fine-tune model, which is called "Answer-Relating Fine-Tune". More details are in our paper.
The following hyperparameters were used during training:
The evaluations of this model as a Candidate Set Generator in CDGP is as follows:
| P@1 | F1@3 | MRR | NDCG@10 |
|---|---|---|---|
| 13.13 | 9.65 | 19.34 | 24.52 |
| Models | CLOTH | DGen |
|---|---|---|
| BERT | cdgp-csg-bert-cloth | cdgp-csg-bert-dgen |
| SciBERT | cdgp-csg-scibert-cloth | cdgp-csg-scibert-dgen |
| RoBERTa | cdgp-csg-roberta-cloth | cdgp-csg-roberta-dgen |
| BART | cdgp-csg-bart-cloth | cdgp-csg-bart-dgen |
fastText: cdgp-ds-fasttext
None