Downloads · 30 days
37
17% of all-time downloads
Helsinki-NLP/opus-mt-caenes-eo
opus-mt-caenes-eo is a translation model from Helsinki-NLP. Use it when you need text moved from one language to another. The card lists the license as cc-by-4.0.
This repository contains a multilingual MarianMT model for (English, Spanish, Catalan) → Esperanto translation.
Downloads · 30 days
37
17% of all-time downloads
All-time downloads
213
Public
Parameters
76.9M
155 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors154 MB · 98%
From the Hugging Face model README
This repository contains a multilingual MarianMT model for (English, Spanish, Catalan) → Esperanto translation.
The model is loaded and used with transformers as:
from transformers import MarianMTModel, MarianTokenizer
import torch
model_name = "Helsinki-NLP/opus-mt-caenes-eo"
device = "cuda" if torch.cuda.is_available() else "cpu"
model = MarianMTModel.from_pretrained(model_name).to(device)
tokenizer = MarianTokenizer.from_pretrained(model_name)
source_texts = [
"Buenos días, qué tal?",
"Bon dia, com estàs?",
"Good morning, how are you?"
]
inputs = tokenizer(source_texts, return_tensors="pt", padding=True, truncation=True)
inputs = {k: v.to(device) for k, v in inputs.items()}
translated_ids = model.generate(inputs["input_ids"])
translated_texts = tokenizer.batch_decode(translated_ids, skip_special_tokens=True)
for src, tgt in zip(source_texts, translated_texts):
print(f"Source: {src} => Translated: {tgt}")
The model was trained using Tatoeba parallel data, with FLORES-200 used as the development set.
Training sentence-pair counts:
| Language Pair | BLEU | ChrF++ |
|---|---|---|
| spa-epo | 16.25 | 49.10 |
| cat-epo | 21.43 | 51.37 |
| eng-epo | 26.42 | 58.23 |