Downloads · 30 days
25
35% of all-time downloads
Emeritus-21/yoruba-codeswitch-diacritics-long-context
yoruba-codeswitch-diacritics-long-context is a machine learning model from Emeritus-21. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
This model restores missing tone marks and underdots (diacritics) in Yorùbá text within code-switched Yorùbá-English contexts. It is explicitly designed and optimized to handle long-context sequences up to 2048 tokens…
Downloads · 30 days
25
35% of all-time downloads
All-time downloads
71
Public
Parameters
300M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 GB · 100%
From the Hugging Face model README
This model restores missing tone marks and underdots (diacritics) in Yorùbá text within code-switched Yorùbá-English contexts. It is explicitly designed and optimized to handle long-context sequences up to 2048 tokens, enabling the accurate restoration of full paragraphs, documents, and extended conversational text without the truncation or coherence loss typical of standard 512-token models.

Figure 1: Phase 3 training and validation loss curves over 5,398 steps. The model achieved a best validation loss of 0.03448 at step 5398, demonstrating stable convergence and successful adaptation to the 2048-token context window without catastrophic forgetting.
To enable the stable training of a 2048-token context window on consumer-grade hardware (NVIDIA RTX 3070 Laptop, 8GB VRAM), this model was trained using a rigorous 3-Phase Curriculum Fine-Tuning strategy. This methodology prevents catastrophic forgetting of foundational orthographic rules while gradually adapting the model's attention mechanisms to longer, more complex dependencies.
| Phase | Dataset Size | Context Length | Primary Objective |
|---|---|---|---|
| Phase 1: Foundation | ~700,000 samples | 256 tokens | Learn core Yorùbá orthographic rules, tone marks, and subdots from high-frequency short sentences. Establish baseline character-level mapping. |
| Phase 2: Bridging | ~70,000 samples | 512 tokens | Adapt to medium-length paragraphs. Learn cross-sentence grammatical dependencies and maintain diacritic consistency across clause boundaries. |
| Phase 3: Long-Context | ~60,000 samples | 2048 tokens | Master document-level coherence, long-range dependencies, and complex code-switching contexts while retaining strict orthographic precision. |
The final phase was trained under the following strict hyperparameter configuration to ensure maximum stability and performance:
A critical methodological requirement for long-context evaluation is preventing silent truncation. If a model is trained on a maximum of 2048 tokens, evaluating it on sequences longer than 2048 tokens will result in the tokenizer silently dropping the end of the sequence, artificially inflating error rates (CER/WER/DER) due to forced deletion errors.
To ensure scientific rigor and accurate metric reporting, the Phase 3 test set was strictly filtered based on true token length:
All reported evaluation metrics below are computed exclusively on this filtered 5,388-sample test set to guarantee that the model was evaluated fairly within its designed operational limits.
All evaluations were conducted on the filtered test set (5,388 samples ≤ 2048 tokens). The following orthogonal metrics were computed to provide a comprehensive view of model performance:
| Decoding Strategy | Samples | CER ↓ | WER ↓ | DER ↓ | WDER ↓ | chrF ↑ | EM ↑ |
|---|---|---|---|---|---|---|---|
| Greedy Search | 5,388 | 1.69% | 4.30% | 8.47% | 5.01% | 95.33 | 2.38% |
| Beam Search (4) | 5,388 | [Pending] | [Pending] | [Pending] | [Pending] | [Pending] | [Pending] |
Note: Beam Search (num_beams=4) evaluation is currently in progress. Final metrics will be updated in this table upon completion.
from transformers import pipeline
# Load the model pipeline
diacritizer = pipeline(
"text2text-generation",
model="Emeritus-21/yoruba-codeswitch-diacritics-long-context",
device=0 # Set to 0 for GPU acceleration, -1 for CPU
)
# Input code-switched text without diacritics
text = "mo ri oko ni ilu Eko because I was driving yesterday"
# Generate restored text (Beam search recommended for maximum accuracy)
result = diacritizer(
text,
max_new_tokens=2048,
num_beams=4,
early_stopping=True
)
print(result[0]['generated_text'])
# Expected output: "mo rí okò ní ìlú Èkó because I was driving yesterday"