Downloads ยท 30 days
453
8% of all-time downloads
Misraj/Baseer__Nakba
Baseer__Nakba is a image-text-to-text model from Misraj. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-nc-sa-4.0.
This repository contains the model weights and inference pipeline for our submission to the NAKBA NLP 2026 Arabic Handwritten Text Recognition (HTR) competition.
Downloads ยท 30 days
453
8% of all-time downloads
All-time downloads
5.5K
Public
Parameters
3.8B
7.5 GB on disk
Likes
6
Public
Click a slice to open those files.
.safetensors7.5 GB ยท 100%
From the Hugging Face model README
This repository contains the model weights and inference pipeline for our submission to the NAKBA NLP 2026 Arabic Handwritten Text Recognition (HTR) competition.
Our approach adapts the 3B-parameter Baseer Vision-Language Model (VLM) to effectively parse and recognize highly cursive, historical Arabic manuscripts. Through a progressive training pipeline, domain-matched data augmentation, and advanced checkpoint merging, this unified model mitigates the challenges of varying writer styles, age-related document degradation, and morphological complexity.
To try our Baseer model for document extraction, please visit: Baseer โ Baseer is the SOTA model on Arabic Document Extraction.
Our final model (Misraj AI) secured 1st place on the official Nakba hidden test set leaderboard.
| Rank | Team | CER | WER |
|---|---|---|---|
| ๐ฅ 1st | Misraj AI | 0.0790 | 0.2440 |
| ๐ฅ 2nd | Oblevit | 0.0925 | 0.3268 |
| ๐ฅ 3rd | 3reeq | 0.0938 | 0.2996 |
| 4th | Latent Narratives | 0.1050 | 0.3106 |
| 5th | Al-Warraq | 0.1142 | 0.3780 |
| 6th | Not Gemma | 0.1217 | 0.3063 |
| 7th | NAMAA-Qari | 0.1950 | 0.5194 |
| 8th | Fahras | 0.2269 | 0.5223 |
| โ | Baseline | 0.3683 | 0.6905 |
Our model was trained using a multi-stage Supervised Fine-Tuning (SFT) curriculum.
All supervised experiments were conducted with standardized hyperparameters across configurations.
| Parameter | Value |
|---|---|
| Hardware | 2ร NVIDIA H100 GPUs |
| Base Model | 3B-parameter Baseer |
| Epochs | 5 |
| Optimizer | AdamW |
| Weight Decay | 0.01 |
| Learning Rate Schedule | Cosine |
| Batch Size | 128 |
| Max Sequence Length | 1200 tokens |
| Input Image Resolution | 644 ร 644 pixels |
| Decoder-Only Learning Rate | 1e-4 |
| Encoder Learning Rate | 9e-6 |
| Decoder Learning Rate (Full Tuning) | 1e-4 |
The model works reliably on images from the Nakba dataset and visually similar historical manuscripts.

This model was merged using the SLERP merge method.
Baseer_Nakba_ep_1Baseer_Nakba_ep_5merge_method: slerp
base_model: Baseer_Nakba_ep_1
models:
- model: Baseer_Nakba_ep_1
- model: Baseer_Nakba_ep_5
parameters:
t:
- value: 0.50
dtype: bfloat16
If you use this model or find our work helpful, please consider citing our paper:
@inproceedings{misrajai2026nakba,
title = {Adapting Vision-Language Models for Historical Arabic Handwritten Text Recognition},
author = {Misraj AI},
booktitle = {Nakba OCR Competition, NLP 2026},
year = {2026}
}