Downloads · 30 days
110
1% of all-time downloads
retrieva-jp/bert-1.3b
bert-1.3b is a fill-mask model from retrieva-jp. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
The RetrievaBERT is the pre-trained Transformer Encoder using Megatron-LM. It is designed for use in Japanese.
Downloads · 30 days
110
1% of all-time downloads
All-time downloads
16K
Public
Parameters
1.3B
5.2 GB on disk
Likes
15
Public
Click a slice to open those files.
.safetensors2.6 GB · 100%
From the Hugging Face model README
The RetrievaBERT is the pre-trained Transformer Encoder using Megatron-LM. It is designed for use in Japanese.
v1.0.1): Bug fix for the model parameters.
The RetrievaBERT is the pre-trained Transformer Encoder using Megatron-LM.
It is designed for use in Japanese.
This model offers several advanced features compared to traditional BERT models:
This model can be used as a Masked Language Model (MLM). However, it is primarily intended to be fine-tuned on downstream tasks. Depending on your use case, follow the appropriate section below.
This model is pre-trained using Masked Language Modeling.
The mask token used is <MASK|LLM-jp>.
Note that you need to set trust_remote_code to True because RetrievaBERT uses a custom model implementation.
Example code for direct use:
from transformers import AutoModelForMaskedLM, AutoTokenizer, pipeline
model_id = "retrieva-jp/bert-1.3b"
model = AutoModelForMaskedLM.from_pretrained(model_id, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(model_id)
pipe = pipeline("fill-mask", model=model, tokenizer=tokenizer)
text = "こんにちは!私の名前は<MASK|LLM-jp>です!"
print(pipe(text))
RetrievaBERT is compatible with Hugging Face's AutoModels. To fine-tune RetrievaBERT for your specific task, use the corresponding AutoModel class. For detailed configuration, refer to the config.json file.
The RetrievaBERT model was pre-trained on the reunion of five datasets:
The model was trained on 180 billion tokens using the above dataset.
The model was trained on 4 to 32 H100 GPUs with a batch size of 1,024. We adopted the curriculum learning which is similar to the Sequence Length Warmup and training with the following sequence lengths and number of steps.
The model was trained on the following hyperparameters.
We fine-tuned the following models and evaluated them on the JGLUE development set. We adjusted the learning rate and training epochs for each model and task in accordance with the JGLUE paper.
| Model | MARC-ja/acc | JSTS/pearson | JSTS/spearman | JNLI/acc | JSQuAD/EM | JSQuAD/F1 | JComQA/acc |
|---|---|---|---|---|---|---|---|
| tohoku-nlp/bert-base-japanese-v3 | 0.957 | 0.914 | 0.876 | 0.906 | 0.878 | 0.946 | 0.849 |
| tohoku-nlp/bert-large-japanese-v2 | 0.959 | 0.916 | 0.877 | 0.901 | 0.884 | 0.951 | 0.867 |
| ku-nlp/deberta-v3-base-japanese | 0.958 | 0.925 | 0.890 | 0.902 | 0.925 | 0.910 | 0.882 |
| retrieva-jp/bert-1.3b | 0.959 | 0.917 | 0.881 | 0.898 | 0.875 | 0.874 | 0.827 |
The RetrievaBERT model is based on BERT with the following hyperparameters:
As mentioned earlier, the main differences from the original BERT are:
This model is based on results obtained from the TSUBAME deep-learning mini-camp.
The model was trained using Megatron-LM.
https://note.com/retrieva/n/n715bea2c2cd1 (in Japanese)
Satoru Katsumata, Daisuke Kimura, Jiro Nishitoba