Downloads · 30 days
0
institutional/institutional-books-visual-elements-orientation
institutional-books-visual-elements-orientation is a image classification model from institutional. Use it when you need a label for an image. The card lists the license as apache-2.0.
A 4‑class image classification model that predicts the rotation correction needed to restore visual elements from digitized book page scans to upright orientation. This model operates on cropped regions (e.g., images,…
Downloads · 30 days
0
Access
Public
Updated Aug 20, 2026
Repo size
213 MB
Likes
0
Public
Click a slice to open those files.
.pth213 MB · 100%
From the Hugging Face model README
A 4‑class image classification model that predicts the rotation correction needed to restore visual elements from digitized book page scans to upright orientation. This model operates on cropped regions (e.g., images, diagrams, ornaments) and is intended as a post-processing stage after visual-element detection.
More information:
See also:
The Institutional Data Initiative at Harvard Law School Library works with knowledge institutions—from libraries and museums to cultural groups and government agencies—to refine and publish their collections as data. Reach out to collaborate on your collections.
The model predicts the inverse rotation required to correct each crop back to upright. Input crops may be in any of four orientations; the model outputs one of four rotation labels.
The labels are “correction actions” to make the image upright:
| Class index | Label | Description |
|---|---|---|
| 0 | upright | No rotation needed (already upright) |
| 1 | rotate_90_clockwise | Rotate 90° clockwise to correct |
| 2 | rotate_180 | Rotate 180° to correct |
| 3 | rotate_90_counterclockwise | Rotate 90° counter-clockwise to correct |
Classes are mutually exclusive.
This model classifies the orientation of already-detected visual elements from digitized book pages.
Primary use cases:
Out of scope:
Source data consists of 7,904 manually curated crops of visual elements from the Institutional Books collection.
Original (pre-synthetic) label distribution (estimated, by source orientation):
To avoid this extreme imbalance and to directly learn the correction operation:
Resulting synthetic orientation dataset:
Each original crop contributes four synthetic samples (one per orientation), producing a balanced label distribution across the four classes in the synthetic set.
| Parameter | Value |
|---|---|
| Backbone | EfficientNetV2-M |
| Classifier head | Dropout(p=0.3) → Linear(1280, 4) |
| Image size (train) | Resize(512×512) → RandomCrop(480×480) |
| Image size (val/test) | Resize(512×512) → CenterCrop(480×480) |
| Batch size | 32 |
| Max epochs | 20 |
| Optimizer / LR | Not specified (standard schedule) |
| Normalization | ImageNet mean/std |
| Hardware | Single NVIDIA GH200 GPU |
| Total training time | 58 min 13 sec (20 epochs) |
| Avg. per epoch | ~2 min 55 sec |
| Train samples/epoch | 25,292 (~791 steps/epoch) |
| Throughput | ~145 images/sec |
Preprocessing:
Resize(512×512)RandomCrop(480×480)ColorJitter(brightness=0.2, contrast=0.2, saturation=0.1)Stochastic augmentations:
| Augmentation | Implementation | Probability |
|---|---|---|
| Random auto-contrast | RandomAutocontrast | p = 0.3 |
| Random invert | RandomInvert | p = 0.15 |
| Random grayscale | RandomGrayscale | p = 0.2 |
| Gaussian blur | GaussianBlur(k=5, σ=0.1–2.0) | p = 1.0 |
| Random erasing | RandomErasing(scale=0.02–0.15) | p = 0.2 |
Validation/Test preprocessing:
Resize(512×512)CenterCrop(480×480)Normalize (ImageNet stats)On the validation set, accuracy increases steadily and plateaus around 90%, while validation loss begins to rise after approximately epoch 10, indicating moderate overfitting. Heavy augmentations successfully limit overfitting enough that the held-out test set slightly outperforms validation.
Final epoch (20):
On the held-out test set (3,164 samples; 791 per class):
Per-class accuracy:
| Class | Accuracy | Correct / Total |
|---|---|---|
| upright | 91.78% | 726 / 791 |
| rotate_90_clockwise | 89.76% | 710 / 791 |
| rotate_180 | 91.91% | 727 / 791 |
| rotate_90_counterclockwise | 91.91% | 727 / 791 |
Additional notes:
upright), 81 (rotate_90_clockwise), 64 (rotate_180), and 64 (rotate_90_counterclockwise).Typical inference settings:
| Parameter | Value |
|---|---|
| Image size | 512×512 resize → 480×480 center crop |
| Batch size | 32 (tune for available GPU memory) |
| Normalization | ImageNet mean/std |
| Output | 4-way softmax over orientation labels |
The top-1 prediction corresponds to the rotation to apply to make the crop upright.
import torch
import torch.nn as nn
import torchvision.models as models
from torchvision import transforms
from huggingface_hub import hf_hub_download
from PIL import Image
# Download weights (the repo ships a state_dict at weights/weights.pth)
model_path = hf_hub_download(
repo_id="institutional/institutional-books-visual-elements-orientation",
filename="weights/weights.pth",
)
# Build the architecture and load the state_dict
model = models.efficientnet_v2_m(weights=None)
num_features = model.classifier[1].in_features
model.classifier = nn.Sequential(
nn.Dropout(p=0.3, inplace=True),
nn.Linear(num_features, 4),
)
state_dict = torch.load(model_path, map_location="cuda", weights_only=True)
model.load_state_dict(state_dict)
model.to("cuda")
model.eval()
# Preprocessing: match validation/test pipeline
preprocess = transforms.Compose([
transforms.Resize((512, 512)),
transforms.CenterCrop(480),
transforms.ToTensor(),
transforms.Normalize(
mean=[0.485, 0.456, 0.406], # ImageNet
std=[0.229, 0.224, 0.225],
),
])
idx_to_label = {
0: "upright",
1: "rotate_90_clockwise",
2: "rotate_180",
3: "rotate_90_counterclockwise",
}
# A high confidence threshold (0.99) is applied: predictions below it
# default to "upright" to minimize false corrections.
CONFIDENCE_THRESHOLD = 0.99
def predict_orientation(path):
img = Image.open(path).convert("RGB")
x = preprocess(img).unsqueeze(0).to("cuda")
with torch.no_grad():
logits = model(x)
probs = torch.softmax(logits, dim=1)[0]
top1 = int(torch.argmax(probs))
conf = float(probs[top1])
label = idx_to_label[top1] if conf >= CONFIDENCE_THRESHOLD else "upright"
return label, conf, probs.cpu().tolist()
label, conf, all_probs = predict_orientation("crop.jpg")
print(f"Predicted correction: {label}, confidence: {conf:.3f}")
from PIL import Image
def apply_correction(img, label):
if label == "upright":
return img
elif label == "rotate_90_clockwise":
return img.rotate(-90, expand=True)
elif label == "rotate_180":
return img.rotate(180, expand=True)
elif label == "rotate_90_counterclockwise":
return img.rotate(90, expand=True)
else:
raise ValueError(f"Unknown label: {label}")
@misc{mendez2026institutionalbooksvisual,
title={Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections},
author={Jimmy Mendez and Matteo Cargnelutti and David Lowry-Duda and Catherine Brobston and Salwa Ismail and Greg Leppert and Amanda Watson and Jonathan Zittrain},
year={2026},
eprint={2608.18957},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.18957},
}