Downloads · 30 days
0
Fallovski/french-serer-transformer-scratch
french-serer-transformer-scratch is a translation model from Fallovski. Use it when you need text moved from one language to another. The card lists the license as other.
Part of a benchmark of six configurations for French→Serer neural machine translation. Serer is a critically low-resource Niger-Congo language (~1.2M speakers, Senegal/Gambia), phylogenetically close to the well-resou…
Downloads · 30 days
0
Access
Public
Updated Aug 14, 2026
Parameters
4.1M
16.7 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.4 MB · 98%
From the Hugging Face model README
Part of a benchmark of six configurations for French→Serer neural machine translation. Serer is a critically low-resource Niger-Congo language (~1.2M speakers, Senegal/Gambia), phylogenetically close to the well-resourced Wolof.
french_serer_transformer_scratch-epoch=23-val_bleu=11.9665.ckptLR: 0.0003NUM_EPOCHS: 30D_MODEL: 128N_ENCODER_LAYERS: 3N_HEADS: 4DROPOUT: 0.3This model has no Atlantic-family prior (its pretraining does not include Wolof or any close relative of Serer), unlike the NLLB-based configurations in this project. In our evaluation, this configuration was judged the most linguistically reliable by a native Serer-speaking expert despite a lower BLEU score than NLLB-based alternatives — see the paper for the full BLEU/quality decorrelation analysis.
Research on French-to-Serer machine translation on a corpus that is ~90% religious (Bible) register, ~10% educational glossaries, primarily Siin dialect. Not validated for legal, medical, emergency, or fully autonomous publication use. Private repository — not intended for public deployment in its current state.
| Metric | Value |
|---|---|
| BLEU (test, beam=5) | 12.0405 |
| chrF | n/a |
| ROUGE-1 | 0.3487 |
| ROUGE-L | 0.2985 |
| BERTScore-F1 | 0.863 |
| Test loss | 2.7447 |
Evaluated on the held-out test split (2890 sentence pairs, SHA-256 of the split:
01d14d982a3c0cce172b5099e2db064d05bdef89677fe72a45f56715d6364ee2). Metrics were computed with the project's own evaluation scripts
(not copied from the manuscript without independent reproduction); the training and
evaluation code is kept in a private repository, available on request.
Parallel corpus of 23113 train / 2889 val / 2890 test French–Serer sentence pairs, built primarily from religious texts (Bible, ~90%) and educational glossaries (~10%), predominantly Siin dialect. Preprocessing: Unicode normalization, exact-duplicate removal, length-ratio filtering (1:3–3:1). Document-level splitting was not possible (no document identifiers available); the split is at the sentence level with a fixed seed. Full provenance, licensing, and consent documentation are kept in a private dataset card, available on request, prior to any public release.