Downloads · 30 days
0
pubmed-ophtha/detection-models
detection-models is a image classification model from pubmed-ophtha. Use it when you need a label for an image. The card lists the license as mit.
This repository contains the three detection and classification models used in the PubMed-Ophtha dataset pipeline for parsing ophthalmological figures from scientific publications.
Downloads · 30 days
0
Access
Public
Updated Sep 7, 2026
Repo size
675 MB
Likes
0
Public
Click a slice to open those files.
.pth675 MB · 100%
From the Hugging Face model README
This repository contains the three detection and classification models used in the PubMed-Ophtha dataset pipeline for parsing ophthalmological figures from scientific publications.
Paper: Hallitschke V.J., Eickhoff C., Berens P. Scientific Domain Knowledge Improves Vision-Language Fundus Models. arXiv:2605.02720 (2026).
| Preprint | arXiv:2605.02720 |
| Dataset | pubmed-ophtha/PubMed-Ophtha |
| Dataset pipeline | berenslab/pubmed-ophtha |
| PDF parser | berenslab/pmo-parser |
| CLIP experiments | berenslab/pmo-experiments |
| Figure-parsing models | This repository |
| PubMed-Ophtha CLIP | PubMed-Ophtha CLIP Models |
| Paper checkpoints | pubmed-ophtha/experiment-checkpoints |
The repository contains three model checkpoints under models/:
| Directory | Checkpoint | Framework | Architecture | Task | Classes |
|---|---|---|---|---|---|
imaging_type_detection_1515892632/ | model_0003909.pth | PyTorch (Detectron2) | RetinaNet + ResNet FPN | Image type detection | CFP, OCT, Retinal Imaging, Other |
panel_detection_1020880423/ | model_0026865.pth | PyTorch (Detectron2) | RetinaNet + ResNet FPN | Panel & identifier detection | Panel, Label |
mark_status_classifier_482239176/ | model_epoch_7.pth | PyTorch | ResNet-50 | Mark status classification | Plain, Annotated |
Each Detectron2 model directory also contains a config.yaml required for inference.
Detects panels and panel identifier labels (e.g. "A", "B") within multi-panel figures. Trained on the PubMed-Ophtha-Annotation dataset merged with PanelSeg and ImageCLEF2016, starting from an ImageCLEF2016-pretrained checkpoint.
Detects individual images within a panel and assigns each a retinal imaging modality: color fundus photography (CFP), optical coherence tomography (OCT), retinal imaging (ultra-wide field / fluorescein angiography), or other (graphs, ultrasound, etc.).
A ResNet-50 binary classifier applied to cropped image regions detected by the image type model. Predicts whether an image contains annotation marks such as arrows, dots, or bounding boxes.
Models are consumed by the pubmed-ophtha
Python package. Download all weights with:
pip install torch==2.8.0 torchvision==0.23.0 setuptools
pip install --no-build-isolation "git+https://github.com/berenslab/[email protected]"
pubmed-ophtha dataset pull-models --local-dir .
Or download directly via huggingface_hub:
from huggingface_hub import snapshot_download
snapshot_download(repo_id="pubmed-ophtha/detection-models", local_dir=".")
After downloading, run inference via the DetectronFigureSplitter:
from pubmed_ophtha.figure_splitting.detectron_figure_splitter import DetectronFigureSplitter
from pubmed_ophtha.const.models import get_default_model_args
splitter = DetectronFigureSplitter(**get_default_model_args())
with open("figure.png", "rb") as f:
image_bytes = f.read()
predictions = splitter.predict(image_bytes)
# Keys: pred_boxes, pred_classes, scores,
# secondary_pred_classes, secondary_scores, keep_after_nms
pred_classes contains Panel/Label detections from the panel detection model followed
by CFP/OCT/Retinal Imaging/Other detections from the image type model.
secondary_pred_classes contains the Plain/Annotated mark status for each image
detection (set to "None" for panel detections).
Both RetinaNet models use a ResNet backbone with FPN, finetuned from an ImageCLEF2016-pretrained Detectron2 checkpoint on the PubMed-Ophtha-Annotation dataset. The ResNet-50 classifier was trained from an ImageNet-pretrained checkpoint for 35 epochs with random cropping, flips, affine transformations, and color augmentation.
The ground-truth annotations used for training are available as part of the PubMed-Ophtha dataset: huggingface.co/datasets/pubmed-ophtha/PubMed-Ophtha
@misc{hallitschke2026scientific,
title={Scientific Domain Knowledge Improves Vision-Language Fundus Models},
author={Verena Jasmin Hallitschke and Carsten Eickhoff and Philipp Berens},
year={2026},
eprint={2605.02720},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.02720},
}
MIT — see LICENSE.