Downloads · 30 days
558
12% of all-time downloads
ChristopherA08/IndoELECTRA
IndoELECTRA is a machine learning model from ChristopherA08. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
ELECTRA is a new method for self-supervised language representation learning. This repository contains the pre-trained Electra Base model (tensorflow 1.15.0) trained in a Large Indonesian corpus (~16GB of raw text | ~…
Downloads · 30 days
558
12% of all-time downloads
All-time downloads
4.7K
Public
Repo size
876 MB
Likes
1
Public
Click a slice to open those files.
.bin438 MB · 100%
From the Hugging Face model README
ELECTRA is a new method for self-supervised language representation learning. This repository contains the pre-trained Electra Base model (tensorflow 1.15.0) trained in a Large Indonesian corpus (~16GB of raw text | ~2B indonesian words). IndoELECTRA is a pre-trained language model based on ELECTRA architecture for the Indonesian Language.
This model is base version which use electra-base config.
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("ChristopherA08/IndoELECTRA")
model = AutoModel.from_pretrained("ChristopherA08/IndoELECTRA")
tokenizer.encode("hai aku mau makan.")
[2, 8078, 1785, 2318, 1946, 18, 4]
The training of the model has been performed using Google's original Tensorflow code on eight core Google Cloud TPU v2. We used a Google Cloud Storage bucket, for persistent storage of training data and models.