Downloads · 30 days
61
73% of all-time downloads
MWirelabs/pnar-mt
pnar-mt is a translation model from MWirelabs. Use it when you need text moved from one language to another. It is set up for transformers. The card lists the license as cc-by-4.0.
This model is a fine-tuned version of facebook/nllb-200-distilled-600M for Pnar (ISO 639-3: pbv) ↔ English machine translation.
Downloads · 30 days
61
73% of all-time downloads
All-time downloads
84
Public
Parameters
615M
2.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.5 GB · 99%
From the Hugging Face model README
This model is a fine-tuned version of
facebook/nllb-200-distilled-600M
for Pnar (ISO 639-3: pbv) ↔ English machine translation.
The model corresponds to the Gold + Silver (Model C) condition described in:
When Silver Isn't Enough: Data Quality Effects on Pnar-English Machine Translation
The model was trained using a combination of a manually translated Gold corpus and automatically generated Silver corpus. The combined dataset produced the strongest overall performance among the evaluated NLLB-200 configurations.
pnar_Latneng_Latn5e-5The Pnar language token was added to the tokenizer and model embedding matrix for fine-tuning.
The Gold corpus contains 4,826 Pnar-English training pairs.
English source sentences were obtained from the Tatoeba corpus and translated into Pnar by two paid native Pnar translators. All translations were human-verified.
The Silver corpus contains 48,519 training pairs derived from the Pnar-English portion of FineTranslations.
The source Pnar text originates primarily from Wyrta, a regional Pnar news media outlet. The English translations were synthetically generated using Gemma 3 27B.
The Silver data therefore primarily represents the news domain.
| Split | Gold | Silver | Combined |
|---|---|---|---|
| Train | 4,826 | 48,519 | 53,345 |
| Validation | 604 | 6,012 | 604 |
| Test | 604 | 5,937 | 604 |
The Gold test set is used as the common evaluation benchmark.
Evaluation was performed on the Gold test set.
| Direction | BLEU | ChrF | TER | COMET |
|---|---|---|---|---|
| English → Pnar | 30.39 | 53.29 | 52.74 | 0.693 |
| Pnar → English | 26.02 | 46.75 | 60.37 | 0.700 |
BLEU, ChrF, TER and COMET were computed using the Hugging Face evaluate library.
COMET should be interpreted as a complementary neural metric because its underlying reference models have not been validated specifically for Pnar.
Human evaluation was conducted on 50 English → Pnar translations from the Gold test set.
Two native Pnar speakers independently evaluated outputs on a 1–5 scale.
| Metric | Score |
|---|---|
| Adequacy | 4.25 / 5 |
| Fluency | 4.53 / 5 |
| Quadratic weighted κ — Adequacy | 0.70 |
| Quadratic weighted κ — Fluency | 0.91 |
The results indicate generally adequate meaning transfer and fluent Pnar output.
This model is intended for:
The model should not be treated as a fully reliable translation system.
The Gold training corpus is relatively small, with only 4,826 training pairs. The Gold and Silver datasets also originate from different domains, meaning that the observed advantage of Gold data may reflect both data quality and domain similarity.
The model may perform poorly on:
Human evaluation found examples where the output was fluent but did not fully preserve the meaning of the English source.
The experiments used a fixed hyperparameter configuration and a single training run per condition. Multi-seed experiments and different Gold/Silver mixing ratios were not evaluated.
Pnar is a low-resource language with limited publicly available NLP resources. Model outputs may contain mistranslations, omissions, hallucinations, or inappropriate lexical choices.
For high-stakes applications, translations should be reviewed by a qualified Pnar speaker.
If you use this model, please cite:
@article{tekcham2026silver,
title={When Silver Isn't Enough: Data Quality Effects on Pnar-English Machine Translation},
author={Tekcham, Riya and Asma, Fitha and Sulfeekhar, Badal Nyalang},
year={2026}
}
This work was supported by MWire Labs, which provided computational resources and support for the research, including assistance with data creation and annotation.