Downloads · 30 days
41
12% of all-time downloads
Ali0044/Qalam_Net_V2
Qalam_Net_V2 is a image-text-to-text model from Ali0044. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
<div align="center" <img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-blue?style=for-the-badge" alt="Hugging Face Model" <img src="https://img.shields.io/badge/Python-3.8%2B-blue?style=for-the…
Downloads · 30 days
41
12% of all-time downloads
All-time downloads
345
Public
Parameters
334M
1.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
Qalam-Net V2 (قلم-نت) is a high-performance Arabic Optical Character Recognition (OCR) system. Built on the TrOCR (Transformer-based OCR) architecture, it achieves superior accuracy by treats OCR as a sequence-to-sequence problem, mapping visual features directly to text tokens.
The model utilizes a Vision-Encoder-Decoder framework, specifically optimized for the complexities of Arabic script (ligatures, cursive nature, and right-to-left orientation).
graph TD
A[Input Arabic Image] --> B[ViT Encoder]
B -->|Visual Embeddings| C[Cross-Attention]
D[Previous Tokens] --> E[RoBERTa Decoder]
E --> C
C --> F[Next Token Prediction]
F -->|Generated Text| G[Final Arabic Transcription]
subgraph "Encoder (Vision Transformer)"
B
end
subgraph "Decoder (Language Model)"
E
end
mssqpi/Arabic-OCR-Dataset for robust handling of various Arabic fonts and styles.microsoft/trocr-base-handwritten.[!IMPORTANT] Qalam-Net V2 differs from traditional OCR by eliminating the need for an external language model or a separate CTC (Connectionist Temporal Classification) layer.
The model was fine-tuned for 1 epoch on a high-quality selection of 5,000 samples.
| Metric | Value |
|---|---|
| Training Samples | 5,000 |
| Optimizer | AdamW |
| Learning Rate | 3e-5 |
| Convergence (Loss) | 9.5 → 0.03 |
[!TIP] Even with a single epoch, the model reached a training loss of 0.03, indicating highly efficient transfer learning from the base TrOCR weights.
pip install transformers datasets Pillow torch
import torch
from PIL import Image, ImageDraw, ImageFont
from transformers import TrOCRProcessor, VisionEncoderDecoderModel
MODEL_NAME = "Ali0044/Qalam_Net_V2"
processor = TrOCRProcessor.from_pretrained(MODEL_NAME)
model = VisionEncoderDecoderModel.from_pretrained(MODEL_NAME)
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)
model.eval()
def run_ocr(image):
pixel_values = processor(image, return_tensors="pt").pixel_values.to(device)
with torch.no_grad():
generated_ids = model.generate(pixel_values)
return processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
image = Image.new('RGB', (200, 50), color = 'white')
d = ImageDraw.Draw(image)
try:
font = ImageFont.truetype("/usr/share/fonts/truetype/dejavu/DejaVuSans-Bold.ttf", 20)
except IOError:
font = ImageFont.load_default()
d.text((10,10), "المتميزة", fill=(0,0,0), font=font)
print(f"Predicted Transcription: {run_ocr(image)}")
image.show()
</details>
Contributions are what make the open-source community an amazing place to learn, inspire, and create.
Ali0044.