Downloads · 30 days
0
VigneshPR/doclayout-yolo-indic
doclayout-yolo-indic is a object detection model from VigneshPR. Use it when you need objects located in an image. It is set up for doclayout-yolo. The card lists the license as apache-2.0.
Real-time document layout detection for 12 Indic scripts, built on DocLayout-YOLO (YOLOv10-m + GL-CRM).
Downloads · 30 days
0
Access
Public
Updated Aug 15, 2026
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.pt1.5 GB · 100%
From the Hugging Face model README
Real-time document layout detection for 12 Indic scripts, built on DocLayout-YOLO (YOLOv10-m + GL-CRM).
This repository accompanies an M.Tech dissertation (Vignesh P, BITS Pilani WILP). It contains four model checkpoints that together form a controlled ablation study — and its central result is an honest negative finding: the two proposed adaptation techniques did not improve over a simple baseline. That finding, proven and explained, is the main scientific contribution.
Document Layout Analysis (DLA) is the first step in digitising any document: before OCR can read the text, a model must find the structure — which regions are paragraphs, headings, tables, figures, lists, and so on. Every downstream stage inherits the errors of this step.
Fast, modern DLA detectors are trained almost entirely on English and Chinese documents and generalise poorly to Indic scripts, because those scripts are visually different:
Over 1.4 billion people use Indic languages, yet no public real-time detector had been demonstrably adapted and evaluated across the major scripts. This project builds one — and rigorously tests whether its own adaptation ideas actually help.
The four files differ only in what training happened before the final fine-tuning on the IndicDLP benchmark. Because everything else is held identical, comparing them isolates the effect of each proposed contribution — this is what licenses causal, not correlational, conclusions.
| File | Training pipeline | Classes | Val mAP@[.5:.95] | Test mAP@[.5:.95] |
|---|---|---|---|---|
config_A_reported_42cls.pt | Public checkpoint → fine-tune. No synthetic pretraining, no self-training. ← the reported model | 42 | 0.379 | 0.364 |
config_B_synthetic_42cls.pt | Synthetic pretrain → fine-tune. Isolates synthetic pretraining. | 42 | 0.353 | 0.252 |
config_C_fullpipeline_42cls.pt | Synthetic → self-train → fine-tune (the full proposed pipeline). | 42 | 0.328 | — |
selftrained_intermediate_9cls.pt | Intermediate self-trained checkpoint, before fine-tuning. | 9 | — | 0.577 (in-domain, Bengali) |
Which to use: for inference, use config_A_reported_42cls.pt — it is the best model and the one the
dissertation reports. (It is a stripped, deployment-ready checkpoint, ~41 MB; the others retain optimizer
state and are ~242 MB.)
text_body, headline, table, figure, caption, advertisement, sidebar, pull-quote, decorative-frame — a coarse, script-agnostic set used for synthetic pretraining and self-training.The detection head is re-shaped between stages (27 → 9 → 42 classes); the backbone and neck weights transfer across all stages, and only the final classifier changes width.
config_A)| Input size | Test mAP@[.5:.95] | Note |
|---|---|---|
| 640 px | 0.329 | fastest; ~90% of peak accuracy at a fraction of the latency |
| 1024 px | 0.364 | peak accuracy (training resolution) |
| 1280 px | 0.353 | strictly dominated — slower and less accurate (train-test resolution mismatch) |
Neither proposed contribution improved over plainly fine-tuning the public checkpoint. Config A (no synthetic, no self-training) is the best model — by +11.2 mAP on the test set over Config B.
This is statistically significant across all 12 scripts:
Mechanism — catastrophic forgetting. The validation→test gap widens from 2.6 to 11.2 mAP under synthetic pretraining — the signature of damaged generalisation, not mere no-gain. The synthetic corpus (23 templates, 9 classes) is narrower than the base model's original DocSynth-300K pretraining; 15 epochs to 0.985 synthetic-validation over-specialised the backbone, and the coarse 9-class intermediate ontology erased fine distinctions the 42-class task then had to relearn from only 12,082 fine-labelled images.
In short: the techniques failed, but the science succeeded — a falsified hypothesis, proven under controlled conditions and mechanistically explained, that saves others from the same mistake.
config_B and config_CThese are provided for transparency and reproducibility of the ablation — they are not recommended for deployment, as they under-perform Config A. They let others verify the negative result independently.
from huggingface_hub import hf_hub_download
from doclayout_yolo import YOLOv10
# download the reported model from this repo
ckpt = hf_hub_download("VigneshPR/doclayout-yolo-indic", "config_A_reported_42cls.pt")
model = YOLOv10(ckpt)
# run on a document page
result = model.predict("page.jpg", imgsz=1024, conf=0.25)[0]
result.plot() # annotated image (numpy array, BGR)
print(len(result.boxes), "regions")
for c in result.boxes.cls.tolist():
print(result.names[int(c)])
Install the runtime:
pip install git+https://github.com/opendatalab/DocLayout-YOLO.git
If you use these models or the finding, please cite the dissertation:
@mastersthesis{vignesh2026doclayoutindic,
title = {DocLayout-YOLO-Indic: Cross-Script Document Layout Analysis for Indic Scripts
using Synthetic Pretraining, Self-Training and Controlled Ablation},
author = {Vignesh P},
school = {BITS Pilani (WILP)},
year = {2026}
}
Built on DocLayout-YOLO (YOLOv10-m + GL-CRM).