Downloads · 30 days
44
11% of all-time downloads
VanModers114/East_Frisian_NLLB_Model
East_Frisian_NLLB_Model is a translation model from VanModers114. Use it when you need text moved from one language to another. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
VanModers114/EastFrisianNLLBModel is the main Oostfräisk Ooversetter model. It is a fine-tuned version of facebook/nllb-200-distilled-600M for translation between East Frisian Low Saxon, German, and English.
Downloads · 30 days
44
11% of all-time downloads
All-time downloads
412
Public
Parameters
615M
34.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt14.8 GB · 60%
From the Hugging Face model README
VanModers114/East_Frisian_NLLB_Model is the main Oostfräisk Ooversetter model. It is a fine-tuned version of facebook/nllb-200-distilled-600M for translation between East Frisian Low Saxon, German, and English.
Supported and trained directions are:
deu_Latn) → East Frisian (frs_Latn)frs_Latn) → German (deu_Latn)eng_Latn) → East Frisian (frs_Latn)frs_Latn) → English (eng_Latn)This is a research-oriented low-resource translation model. It is not a certified translation service and should not be relied upon without human review for consequential decisions.
| Property | Value |
|---|---|
| Developer | Oostfräisk Instituut, VanModers114 |
| Model type | Encoder-decoder Transformer (M2M100ForConditionalGeneration) |
| Base model | facebook/nllb-200-distilled-600M |
| Parameters | NLLB distilled 600M class |
| Context limit | 1,024 positions; fine-tuning truncates to 512 tokens |
| East Frisian language token | frs_Latn |
| License | CC BY-NC 4.0, inherited from the NLLB base model |
| Source code | VanModers/oostfraeisk_ooversetter |
NLLB does not include a dedicated East Frisian token. Fine-tuning adds frs_Latn in an unused embedding slot and initializes that embedding from Dutch nld_Latn. German and English retain their native NLLB tokens.
The final model from the recorded run was evaluated German → East Frisian on a separate 100-pair project validation corpus:
| Metric | Result |
|---|---|
| SacreBLEU | 62.06 |
| Teacher-forced token accuracy | 0.8937 |
| Average loss | 0.8808 |
Important qualifications:
The model uses NLLB-style source and target language control. The compatibility setup below also handles the strict boolean/integer configuration validation present in transformers 5.4.0:
import json
import torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer, M2M100Config
MODEL_ID = "VanModers114/East_Frisian_NLLB_Model"
# Compatibility for strict transformers 5.4.0 validation of scale_embedding.
scale_field = getattr(M2M100Config, "__dataclass_fields__", {}).get("scale_embedding")
if scale_field is not None and isinstance(scale_field.default, bool):
scale_field.default = int(scale_field.default)
config_path = hf_hub_download(repo_id=MODEL_ID, filename="config.json")
with open(config_path, encoding="utf-8") as config_file:
config_dict = json.load(config_file)
if isinstance(config_dict.get("scale_embedding"), bool):
config_dict["scale_embedding"] = int(config_dict["scale_embedding"])
config = M2M100Config.from_dict(config_dict)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, config=config)
model = AutoModelForSeq2SeqLM.from_pretrained(MODEL_ID, config=config)
model.eval()
# Supported codes: deu_Latn, eng_Latn, frs_Latn
source_language = "deu_Latn"
target_language = "frs_Latn"
text = "Das ist ein Beispielsatz."
frs_id = tokenizer.convert_tokens_to_ids("frs_Latn")
if frs_id == tokenizer.unk_token_id:
raise RuntimeError("The tokenizer does not contain the frs_Latn token")
# Compatibility for tokenizer versions that keep explicit language-code maps.
if hasattr(tokenizer, "lang_code_to_id"):
tokenizer.lang_code_to_id["frs_Latn"] = frs_id
if hasattr(tokenizer, "id_to_lang_code"):
tokenizer.id_to_lang_code[frs_id] = "frs_Latn"
tokenizer.src_lang = source_language
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
target_id = tokenizer.convert_tokens_to_ids(target_language)
with torch.inference_mode():
output = model.generate(
**inputs,
forced_bos_token_id=target_id,
max_new_tokens=512,
num_beams=4,
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
The source repository's nllb_model/app.py contains the maintained Gradio implementation of the same loading approach.
This model is distributed under CC BY-NC 4.0, inherited from facebook/nllb-200-distilled-600M. Review the base model card and full license before use.
If you use the underlying NLLB architecture or weights, cite:
@article{nllb2022,
title = {No Language Left Behind: Scaling Human-Centered Machine Translation},
author = {{NLLB Team} and Costa-jussà, Marta R. and others},
journal = {arXiv preprint arXiv:2207.04672},
year = {2022}
}
Project model and training code: https://github.com/VanModers/oostfraeisk_ooversetter
Open an issue in the source repository or contact the model publisher through Hugging Face or by contacting the Oostfräisk Instituut.