Downloads · 30 days
28
0% of all-time downloads
sarahlintang/IndoBERT
IndoBERT is a machine learning model from sarahlintang. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
IndoBERT is a pre-trained language model based on BERT architecture for the Indonesian Language.
Downloads · 30 days
28
0% of all-time downloads
All-time downloads
6.5K
Public
Repo size
2.7 GB
Likes
2
Public
Click a slice to open those files.
.data-00000-of-000011.4 GB · 60%
From the Hugging Face model README
IndoBERT is a pre-trained language model based on BERT architecture for the Indonesian Language.
This model is base-uncased version which use bert-base config.
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("sarahlintang/IndoBERT")
model = AutoModel.from_pretrained("sarahlintang/IndoBERT")
tokenizer.encode("hai aku mau makan.")
[2, 8078, 1785, 2318, 1946, 18, 4]
This model was pre-trained on 16 GB of raw text ~2 B words from Oscar Corpus (https://oscar-corpus.com/).
This model is equal to bert-base model which has 32,000 vocabulary size.
The training of the model has been performed using Google’s original Tensorflow code on eight core Google Cloud TPU v2. We used a Google Cloud Storage bucket, for persistent storage of training data and models.
We evaluate this model on three Indonesian NLP downstream task: