Downloads · 30 days
0
AlanaBusu/DocEng_Tessarect_mlt
DocEng_Tessarect_mlt is a machine learning model from AlanaBusu. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A fine-tuned Tesseract OCR model for Maltese-language news articles.
Downloads · 30 days
0
Access
Public
Updated Jun 30, 2026
Repo size
3.2 MB
Likes
1
Public
Click a slice to open those files.
.traineddata3.2 MB · 100%
From the Hugging Face model README
A fine-tuned Tesseract OCR model for Maltese-language news articles.
This model was fine-tuned on a synthetic dataset of ~20k PNG image "snippets," each depicting a short excerpt of Maltese text paired with its ground-truth transcription. Training ran for 2000 iterations on top of the base Tesseract Maltese trained data.
The synthetic training images were produced in two stages:
Text generation — Maltese news-style text snippets were generated by querying Gemini 3.1 Flash-Lite to generate 120 pdfs covering 10 different topics. The snippets were based on the character set provided in the competition assets and generated predominately in Maltese with some English covering different fonts, text and background colours.
Snippet Extraction - The snippet extraction process scans each individual PDF, identifies text coordinates, and crops specific areas into individual images. The resulting images are saved alongside a CSV index that links each snippet to its corresponding text, and word count. Gaussian blur is applied to 10% of these snippets.
Manual Noising — Randomised manual degradation was implemented.
Download the trained data file from the Hugging Face Hub and point Tesseract at its containing directory:
import os
from huggingface_hub import hf_hub_download
traineddata_path = hf_hub_download(
repo_id="AlanaBusu/DocEng_Tessarect_mlt",
filename="finetune_tess_extended_2000.traineddata",
)
tessdata_dir = os.path.dirname(traineddata_path)