Downloads · 30 days
24
39% of all-time downloads
harness-race/control-r3
control-r3 is a object detection model from harness-race. Use it when you need objects located in an image. It is set up for transformers. The card lists the license as apache-2.0.
This model is a fine-tuned DetrForObjectDetection (DETR with a ResNet-50 backbone, starting from facebook/detr-resnet-50) tuned to localize layout regions in historical and modern newspaper pages. It predicts bounding…
Downloads · 30 days
24
39% of all-time downloads
All-time downloads
62
Public
Parameters
41.6M
167 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors167 MB · 100%
From the Hugging Face model README
This model is a fine-tuned DetrForObjectDetection (DETR with a ResNet-50 backbone,
starting from facebook/detr-resnet-50)
tuned to localize layout regions in historical and modern newspaper pages.
It predicts bounding boxes for the 7 classes of the
BigLam “Locating Objects Beyond Words” dataset:
| class id | label |
|---|---|
| 0 | Photograph |
| 1 | Illustration |
| 2 | Map |
| 3 | Comics/Cartoon |
| 4 | Editorial Cartoon |
| 5 | Headline |
| 6 | Advertisement |
The base model facebook/detr-resnet-50 is
released under the Apache 2.0 license, so this fine-tune can be freely shared and used.
timm), pretrained on ImageNet.600 × 600 after a smallest-max-size resize.no object class).biglam/loc_beyond_words.The model is intended for document / newspaper-page layout analysis: given a scan or a page image, it detects coarse layout regions such as headlines, photographs/illustrations, advertisements, maps and comics. It is designed as a layout-region detector and is not meant for fine-grained text recognition or OCR (use OCR/HTR systems for reading text).
Example inference:
from transformers import pipeline
detector = pipeline("object-detection", model="harness-race/control-r3")
results = detector("path/to/newspaper_page.png")
# results: list of {label, score, box: {xmin, ymin, xmax, ymax}}
biglam/loc_beyond_words
(BigLam “Locating Objects Beyond Words”, a Library-of-Congress-derived newspaper layout dataset).[x, y, width, height] in pixel coordinates.harness-race/loc_beyond_words_coco;
no annotations were modified.The model was fine-tuned end-to-end (all weights trainable) with the Hugging Face Trainer-style loop
on a single NVIDIA T4 GPU (fp16 AMP), with the RGB images resized and padded to 600 × 600 and light
augmentation (horizontal flip, random brightness/contrast, hue/saturation, random crop with box clipping).
Evaluation is run at the end of every epoch and the checkpoint with the best validation mAP is kept.
Reported on the validation split (712 images), using COCO-style metrics
(torchmetrics.MeanAveragePrecision, box_format=xyxy). Metrics are in %:
| Metric | Value |
|---|---|
| mAP (IoU .5:.95) | 24.65 |
| mAP @ IoU 0.50 | 34.76 |
| mAP @ IoU 0.75 | 28.39 |
| mAR@100 | 37.94 |
Per-class mAP (IoU .5:.95):
| Class | mAP |
|---|---|
| Photograph | 39.02 |
| Illustration | 1.17 |
| Map | 0.03 |
| Comics/Cartoon | 13.21 |
| Editorial Cartoon | 0.00 |
| Headline | 59.51 |
| Advertisement | 59.64 |
The model detects Headline and Advertisement regions very well (>59 mAP) and detects Photograph and Comics/Cartoon reasonably. The rare classes (Illustration, Map, Editorial Cartoon) show very low mAP, which is largely a consequence of heavy class imbalance in the dataset (e.g. only ~215 Map and ~293 Editorial Cartoon object instances across the whole dataset vs ~27.9k Headline instances). More data or class-balancing/oversampling for those classes would improve them.
The raw per-epoch metrics are stored in val_metrics.json in this repository.
DetrForObjectDetection, 100 queries, 6 encoder + 6 decoder layers, d_model=256.config.json, model.safetensors, preprocessor_config.json
(DetrImageProcessorFast, size=600), val_metrics.json, train_detr.py (training script).Based on the DETR model (Carion et al., 2020) and the Transformers library. Dataset from BigLam and the Library of Congress newspaper collections.