Downloads · 30 days
17
5% of all-time downloads
fcrescio/rotdet
rotdet is a machine learning model from fcrescio. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
Downloads · 30 days
17
5% of all-time downloads
All-time downloads
330
Public
Parameters
276K
2.3 MB on disk
Likes
2
Public
Click a slice to open those files.
.png1.2 MB · 52%
From the Hugging Face model README

This is a SimpleCNN model designed to detect whether an image containing text (e.g. a scan of a document) is correctly oriented or rotated. It takes grayscale images as input, resizes them to 128x128 pixels, and outputs a prediction indicating whether the image is "Rotated" or "Normal."
The model consists of three convolutional layers followed by max-pooling operations. These convolutional layers extract features from the input image. The feature maps are then flattened and passed through two fully connected layers, where the final layer outputs a prediction between two classes:
Class 0 (Normal): The image is correctly oriented. Class 1 (Rotated): The image is rotated and needs adjustment. The model is trained on a dataset of rotated and correctly oriented grayscale images. It is capable of accurately distinguishing between the two classes and can be used in applications that involve automatic image processing or document scanning.
Inference To use this model for inference, you can load it using Hugging Face's from_pretrained functionality and pass in an image for orientation prediction.
import torch
import torch.nn as nn
import torch.nn.functional as F
from safetensors.torch import load_file
from PIL import Image
import numpy as np
# Define the corrected SimpleCNN architecture
class SimpleCNN(nn.Module):
def __init__(self):
super(SimpleCNN, self).__init__()
self.conv1 = nn.Conv2d(1, 16, kernel_size=3, stride=1, padding=1) # Adjusted to 16 output channels
self.conv2 = nn.Conv2d(16, 32, kernel_size=3, stride=1, padding=1) # Adjusted to 32 output channels
self.conv3 = nn.Conv2d(32, 32, kernel_size=3, stride=1, padding=1) # Adjusted to 32 output channels
self.pool = nn.MaxPool2d(kernel_size=2, stride=2)
self.fc1 = nn.Linear(32 * 16 * 16, 32) # Adjusted input and output dimensions
self.fc2 = nn.Linear(32, 2) # Adjusted input dimension
def forward(self, x):
x = self.pool(F.relu(self.conv1(x)))
x = self.pool(F.relu(self.conv2(x)))
x = self.pool(F.relu(self.conv3(x)))
x = x.view(x.size(0), -1) # Flatten
x = F.relu(self.fc1(x))
x = self.fc2(x)
return x
# Load the model
model = SimpleCNN()
state_dict = load_file("model.safetensors")
model.load_state_dict(state_dict)
model.eval()
# Function to predict orientation
def predict_orientation(image_path, model):
img = Image.open(image_path).convert('L') # Load image in grayscale
img = img.resize((128, 128)) # Resize to 128x128
img_tensor = torch.tensor(np.array(img) / 255.0, dtype=torch.float32).unsqueeze(0).unsqueeze(0)
with torch.no_grad():
output = model(img_tensor)
is_rotated = torch.argmax(output, dim=1).item() == 1
return "Rotated" if is_rotated else "Normal"
# Example usage
result = predict_orientation("example.jpg", model)
print(f"Image Orientation: {result}")
(HT: https://huggingface.co/khasinski)
The model was trained using standard binary cross-entropy loss and an Adam optimizer. It was trained on grayscale images resized to 128x128 pixels.
Training scripts and evaluation scripts are available at https://github.com/fcrescio/rotdet
The model performs well in scenarios where images need to be automatically detected for correct orientation. However, the performance can vary based on the image quality, input resolution, and types of rotations present in the dataset.
The model is trained only on 90-degree rotations, meaning performance might degrade with other types of rotations (e.g., slight tilts or partial rotations). It is designed to work on grayscale images, so it might not perform optimally on colored or highly textured images. Intended Use The primary use case for this model is in scenarios where the orientation of images needs to be detected or corrected, such as:
The model does not process or output sensitive information. However, users should be aware of potential biases that could be introduced by the training dataset (e.g., specific types of images or orientations might be overrepresented).
If you use this model, please cite the following:
@misc{simplecnn_orientation,
author = {Francesco Crescioli},
title = {SimpleCNN for Image Orientation Detection},
year = {2024},
howpublished = {\url{https://huggingface.co/fcrescio/rotdet}},
}
This model is licensed under the Creative Commons Attribution 4.0 International (CC-BY-4.0) License.
This model was trained using the Docmatix database, which is licensed under the MIT license. As such, the following MIT license applies to the data used in training this model: