Downloads · 30 days
64
14% of all-time downloads
PartiallyTyped/answerable_tydiqa_lm_pretrained_japanese
answerable_tydiqa_lm_pretrained_japanese is a feature extraction model from PartiallyTyped. Use it when you need embeddings to search or compare text. It is set up for transformers.
This is a pretrained model based on rinna/japanese-gpt2-small that has been trained on copenlu/answerabletydiqa, specifically the text field of the Japanese samples for 2 epochs.
Downloads · 30 days
64
14% of all-time downloads
All-time downloads
448
Public
Repo size
909 MB
Likes
0
Public
Click a slice to open those files.
.bin454 MB · 100%
From the Hugging Face model README
This is a pretrained model based on rinna/japanese-gpt2-small that has been trained on copenlu/answerable_tydiqa, specifically the text field of the Japanese samples for 2 epochs.
To use the pretrained head, use: AutoModelForCausalLM.from_pretrained.
from transformers import AutoModelForCausalLM, T5Tokenizer
model_path = "PartiallyTyped/answerable_tydiqa_lm_pretrained_japanese"
model = AutoModelForCausalLM.from_pretrained(path)
tokenizer = T5Tokenizer.from_pretrained(path)
tokenizer.do_lower_case = True # due to some bug of tokenizer config loading