Downloads · 30 days
4.4K
0% of all-time downloads
staka/fugumt-en-ja
fugumt-en-ja is a translation model from staka. Use it when you need text moved from one language to another. It is set up for transformers. The card lists the license as cc-by-sa-4.0.
This is a translation model using Marian-NMT. For more details, please see my repository.
Downloads · 30 days
4.4K
0% of all-time downloads
All-time downloads
1.1M
Public
Repo size
550 MB
Likes
54
Public
Click a slice to open those files.
.bin121 MB · 98%
From the Hugging Face model README
This is a translation model using Marian-NMT. For more details, please see my repository.
This model uses transformers and sentencepiece.
!pip install transformers sentencepiece
You can use this model directly with a pipeline:
from transformers import pipeline
fugu_translator = pipeline('translation', model='staka/fugumt-en-ja')
fugu_translator('This is a cat.')
If you want to translate multiple sentences, we recommend using pySBD.
!pip install transformers sentencepiece pysbd
import pysbd
seg_en = pysbd.Segmenter(language="en", clean=False)
from transformers import pipeline
fugu_translator = pipeline('translation', model='staka/fugumt-en-ja')
txt = 'This is a cat. It is very cute.'
print(fugu_translator(seg_en.segment(txt)))
The results of the evaluation using tatoeba(randomly selected 500 sentences) are as follows:
| source | target | BLEU(*1) |
|---|---|---|
| en | ja | 32.7 |
(*1) sacrebleu --tokenize ja-mecab