Downloads · 30 days
0
yukieos/grocery_classification
grocery_classification is a image classification model from yukieos. Use it when you need a label for an image. The card lists the license as apache-2.0.
This modelcard aims to be a base template for new models. It has been generated using this raw template.
Downloads · 30 days
0
Access
Public
Updated May 30, 2025
Repo size
6.4 MB
Likes
0
Public
Click a slice to open those files.
.pth6.4 MB · 100%
From the Hugging Face model README
This modelcard aims to be a base template for new models. It has been generated using this raw template.
Call the infer_category function to get a grocery label from an image, using OCR first and falling back to a MobileNetV3‐based classifier: category, ocr_text, method = infer_category(image_path) #can replace "image_path" with an address of image query = ocr_text if method=='ocr' else category if method=='manual': query = input("Please type the product you want to search:") print(f"Result: {category} (via {method})")
Non-grocery products (e.g. electronics, clothing)
Packaging text in languages other than English
Highly distorted or tiny product labels
Class bias: only predicts the ~20 grocery categories seen in training
OCR bias: EasyOCR may misread stylized fonts, low-contrast text, or cluttered backgrounds
Failure mode: when OCR and classifier both low-confidence, returns None
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
Use the code below to get started with the model:
pip install torch torchvision easyocr opencv-python pillow
from infer import infer_category
label, raw_text, method = infer_category("image.jpg")
print(f"Label: {label}, OCR saw: {raw_text}, Method: {method}")
~2,000 images collected in local grocery stores, evenly split across 20 product categories (e.g. “apple,” “banana,” “milk,” …). Manually labeled and organized into train/val/test splits.
Resize to 224×224 Random horizontal flips, random rotation ±15° Normalize with ImageNet means/stds
400 held-out images (20 per class) from the same data distribution.
Varying lighting Text printed vs. handwritten Partial occlusion
Overall pipeline accuracy on test set: 89.0%
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Backbone: MobileNetV3-Small, final fc → 20 classes
OCR: EasyOCR reader (en) for text detection + recognition
Loss: CrossEntropy for classification
Training environment: Colab notebook with single GPU
Inference environment: any machine with PyTorch + EasyOCR
NVIDIA T4 (training & eval)
Python 3.10
PyTorch 1.13, torchvision 0.14
EasyOCR 1.4, OpenCV 4.x, Pillow 9.x