Downloads · 30 days
11
3% of all-time downloads
rmihaylov/roberta2roberta-shared-nmt-bg
roberta2roberta-shared-nmt-bg is a machine learning model from rmihaylov. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
This model was introduced in this paper.
Downloads · 30 days
11
3% of all-time downloads
All-time downloads
417
Public
Repo size
2.5 GB
Likes
0
Public
Click a slice to open those files.
.bin1.2 GB · 100%
From the Hugging Face model README
This model was introduced in this paper.
The training data is private English-Bulgarian parallel data.
You can use the raw model for translation from English to Bulgarian.
Here is how to use this model in PyTorch:
>>> from transformers import EncoderDecoderModel, XLMRobertaTokenizer
>>>
>>> model_id = "rmihaylov/roberta2roberta-shared-nmt-bg"
>>> model = EncoderDecoderModel.from_pretrained(model_id)
>>> model.encoder.pooler = None
>>> tokenizer = XLMRobertaTokenizer.from_pretrained(model_id)
>>>
>>> text = """
Others were photographed ransacking the building, smiling while posing with congressional items such as House Speaker Nancy Pelosi's lectern or at her staffer's desk, or publicly bragged about the crowd's violent and destructive joyride.
"""
>>>
>>> inputs = tokenizer.encode_plus(text, max_length=100, return_tensors='pt', truncation=True)
>>>
>>> translation = model.generate(**inputs,
>>> max_length=100,
>>> num_beams=4,
>>> do_sample=True,
>>> num_return_sequences=1,
>>> top_p=0.95,
>>> decoder_start_token_id=tokenizer.bos_token_id)
>>>
>>> print([tokenizer.decode(g.tolist(), skip_special_tokens=True) for g in translation])
['Други бяха заснети да бягат из сградата, усмихвайки се, докато се представят с конгресни предмети, като например лекцията на председателя на парламента Нанси Пелози или на бюрото на нейния служител, или публично се хвалят за насилието и разрушителната радост на тълпата.']