Downloads · 30 days
41
1% of all-time downloads
RUCAIBox/Erya
Erya is a translation model from RUCAIBox. Use it when you need text moved from one language to another. It is set up for transformers. The card lists the license as apache-2.0.
Erya is a pretrained model specifically designed for translating Ancient Chinese into Modern Chinese. It utilizes an Encoder-Decoder architecture and has been trained using a combination of DMLM (Dual Masked Language…
Downloads · 30 days
41
1% of all-time downloads
All-time downloads
3.9K
Public
Repo size
3.2 GB
Likes
10
Public
Click a slice to open those files.
.bin1.1 GB · 100%
From the Hugging Face model README
Erya is a pretrained model specifically designed for translating Ancient Chinese into Modern Chinese. It utilizes an Encoder-Decoder architecture and has been trained using a combination of DMLM (Dual Masked Language Model) and DAS (Disyllabic Aligned Substitution) techniques on datasets comprising both Ancient Chinese and Modern Chinese texts. The detailed information of our work can be found here: RUCAIBox/Erya (github.com)
More information about Erya dataset can be found here: RUCAIBox/Erya-dataset · Datasets at Hugging Face, which can be used to tune the Erya model further for a better translation performance.
>>> from transformers import BertTokenizer, CPTForConditionalGeneration
>>> tokenizer = BertTokenizer.from_pretrained("RUCAIBox/Erya")
>>> model = CPTForConditionalGeneration.from_pretrained("RUCAIBox/Erya")
>>> input_ids = tokenizer("安世字子孺,少以父任为郎。", return_tensors='pt')
>>> input_ids.pop("token_type_ids")
>>> pred_ids = model.generate(max_new_tokens=256, **input_ids)
>>> print(tokenizer.batch_decode(pred_ids, skip_special_tokens=True))
['安 世 字 子 孺 , 年 轻 时 因 父 任 郎 官 。']