Downloads · 30 days
32.6K
2% of all-time downloads
allenai/scibert_scivocab_cased
scibert_scivocab_cased is a machine learning model from allenai. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
This is the pretrained model presented in SciBERT: A Pretrained Language Model for Scientific Text, which is a BERT model trained on scientific text.
Downloads · 30 days
32.6K
2% of all-time downloads
All-time downloads
1.7M
Public
Repo size
1.3 GB
Likes
18
Public
Click a slice to open those files.
.bin442 MB · 50%
From the Hugging Face model README
This is the pretrained model presented in SciBERT: A Pretrained Language Model for Scientific Text, which is a BERT model trained on scientific text.
The training corpus was papers taken from Semantic Scholar. Corpus size is 1.14M papers, 3.1B tokens. We use the full text of the papers in training, not just abstracts.
SciBERT has its own wordpiece vocabulary (scivocab) that's built to best match the training corpus. We trained cased and uncased versions.
Available models include:
scibert_scivocab_casedscibert_scivocab_uncasedThe original repo can be found here.
If using these models, please cite the following paper:
@inproceedings{beltagy-etal-2019-scibert,
title = "SciBERT: A Pretrained Language Model for Scientific Text",
author = "Beltagy, Iz and Lo, Kyle and Cohan, Arman",
booktitle = "EMNLP",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://www.aclweb.org/anthology/D19-1371"
}