Downloads · 30 days
2.5K
3% of all-time downloads
SIKU-BERT/sikubert
sikubert is a fill-mask model from SIKU-BERT. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
Digital humanities research needs the support of large-scale corpus and high-performance ancient Chinese natural language processing tools. The pre-training language model has greatly improved the accuracy of text min…
Downloads · 30 days
2.5K
3% of all-time downloads
All-time downloads
82.4K
Public
Repo size
1.3 GB
Likes
16
Public
Click a slice to open those files.
.bin438 MB · 100%
From the Hugging Face model README
Digital humanities research needs the support of large-scale corpus and high-performance ancient Chinese natural language processing tools. The pre-training language model has greatly improved the accuracy of text mining in English and modern Chinese texts. At present, there is an urgent need for a pre-training model specifically for the automatic processing of ancient texts. We used the verified high-quality “Siku Quanshu” full-text corpus as the training set, based on the BERT deep language model architecture, we constructed the SikuBERT and SikuRoBERTa pre-training language models for intelligent processing tasks of ancient Chinese.
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("SIKU-BERT/sikubert")
model = AutoModel.from_pretrained("SIKU-BERT/sikubert")
We are from Nanjing Agricultural University.