Downloads · 30 days
39
0% of all-time downloads
Ayaka/bart-base-cantonese
bart-base-cantonese is a fill-mask model from Ayaka. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as other.
This is the Cantonese model of BART base. It is obtained by a second-stage pre-training on the LIHKG dataset based on the fnlp/bart-base-chinese model.
Downloads · 30 days
39
0% of all-time downloads
All-time downloads
18.7K
Public
Repo size
2.2 GB
Likes
9
Public
Click a slice to open those files.
.bin439 MB · 48%
From the Hugging Face model README
This is the Cantonese model of BART base. It is obtained by a second-stage pre-training on the LIHKG dataset based on the fnlp/bart-base-chinese model.
This project is supported by Cloud TPUs from Google's TPU Research Cloud (TRC).
Note: To avoid any copyright issues, please do not use this model for any purpose.
from transformers import BertTokenizer, BartForConditionalGeneration, Text2TextGenerationPipeline
tokenizer = BertTokenizer.from_pretrained('Ayaka/bart-base-cantonese')
model = BartForConditionalGeneration.from_pretrained('Ayaka/bart-base-cantonese')
text2text_generator = Text2TextGenerationPipeline(model, tokenizer)
output = text2text_generator('聽日就要返香港,我激動到[MASK]唔着', max_length=50, do_sample=False)
print(output[0]['generated_text'].replace(' ', ''))
# output: 聽日就要返香港,我激動到瞓唔着
Note: Please use the BertTokenizer for the model vocabulary. DO NOT use the original BartTokenizer.
WandB link: 1j7zs802