Downloads · 30 days
14
1% of all-time downloads
HillZhang/pseudo_native_bart_CGEC_thesis
pseudo_native_bart_CGEC_thesis is a machine learning model from HillZhang. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This model is a cutting-edge CGEC model based on Chinese BART-large. It is trained with about 100M pseudo native speaker CGEC training data generated by heuristic rules and human-annotated training data for the thesis…
Downloads · 30 days
14
1% of all-time downloads
All-time downloads
1.3K
Public
Repo size
3 GB
Likes
0
Public
Click a slice to open those files.
.bin1.5 GB · 100%
From the Hugging Face model README
This model is a cutting-edge CGEC model based on Chinese BART-large. It is trained with about 100M pseudo native speaker CGEC training data generated by heuristic rules and human-annotated training data for the thesis domain. More details can be found in our Github and the paper.
pip install transformers
from transformers import BertTokenizer, BartForConditionalGeneration, Text2TextGenerationPipeline
tokenizer = BertTokenizer.from_pretrained("HillZhang/pseudo_native_bart_CGEC_thesis")
model = BartForConditionalGeneration.from_pretrained("HillZhang/pseudo_native_bart_CGEC_thesis")
encoded_input = tokenizer(["北京是中国的都。", "他说:”我最爱的运动是打蓝球“", "我每天大约喝5次水左右。", "今天,我非常开开心。"], return_tensors="pt", padding=True, truncation=True)
if "token_type_ids" in encoded_input:
del encoded_input["token_type_ids"]
output = model.generate(**encoded_input)
print(tokenizer.batch_decode(output, skip_special_tokens=True))
@inproceedings{zhang-etal-2023-nasgec,
title = "{Na}{SGEC}: a Multi-Domain Chinese Grammatical Error Correction Dataset from Native Speaker Texts",
author = "Zhang, Yue and
Zhang, Bo and
Jiang, Haochen and
Li, Zhenghua and
Li, Chen and
Huang, Fei and
Zhang, Min"
booktitle = "Findings of ACL",
year = "2023"
}