Downloads · 30 days
0
hur03/capturemate-image-classifier-v1-5class
capturemate-image-classifier-v1-5class is a image classification model from hur03. Use it when you need a label for an image. It is set up for transformers.
CLIP-based screenshot image classifier for CaptureMate.
Downloads · 30 days
0
Access
Public
Updated Aug 7, 2026
Repo size
352 MB
Likes
0
Public
Click a slice to open those files.
.pt352 MB · 100%
From the Hugging Face model README
CLIP-based screenshot image classifier for CaptureMate.
This model classifies screenshot images into one of five CaptureMate categories:
scheduleshoppingplacememounknownThe model is designed to complement the CaptureMate OCR text classifier by providing visual information when screenshot text alone is ambiguous or incomplete.
Note: This model uses a custom PyTorch classification head on top of a CLIP vision encoder. Therefore, it cannot be loaded directly with
AutoModelForImageClassificationor the standard Hugging Faceimage-classificationpipeline.
This model is a CLIP-based image classification model developed for CaptureMate, an iOS screenshot organization and action recommendation app.
It receives a screenshot image and predicts the most relevant CaptureMate category based on visual information.
openai/clip-vit-base-patch32This repository contains v1 of the 5-class CaptureMate screenshot image classifier.
The model predicts the following categories:
scheduleshoppingplacememounknownFuture versions may improve image-only classification performance as additional screenshot data becomes available.
| ID | Label | Description |
|---|---|---|
| 0 | schedule | Schedule, reservation, ticket, event, or date-related screenshots |
| 1 | shopping | Shopping, product, price, payment, or commerce-related screenshots |
| 2 | place | Place, map, restaurant, store, travel, or location-related screenshots |
| 3 | memo | Text, article, note, content, or general information screenshots |
| 4 | unknown | Ambiguous, image-heavy, low-information, or non-actionable screenshots |
This model can be used to classify screenshot images into the five CaptureMate categories.
Unlike the CaptureMate OCR text classifier, this model analyzes the visual information contained in the screenshot.
The model is especially useful when:
This model is primarily designed to be used together with the CaptureMate OCR text classifier.
The overall CaptureMate classification pipeline uses both OCR text and screenshot image information:
Screenshot
|
+----------+----------+
| |
v v
OCR Extraction Image Analysis
| |
v v
Text Classifier Image Classifier
| |
+----------+----------+
|
v
Score Fusion
|
v
Category Decision
|
v
Action Recommendation
In the current CaptureMate multimodal pipeline, the prediction scores are combined using fixed weighted fusion:
Text weight = 0.7
Image weight = 0.3
Conceptually:
final_score =
0.7 × text_score
+
0.3 × image_score
The category with the highest fused score is selected as the final category prediction.
The image classifier is therefore intended to provide a secondary visual signal rather than replace the OCR text classifier.
This model is not designed to:
Install the required libraries:
pip install torch transformers huggingface_hub pillow
The model files can be downloaded directly from Hugging Face Hub using hf_hub_download.
from huggingface_hub import hf_hub_download
model_path = hf_hub_download(
repo_id="hur03/capturemate-image-classifier-v1-5class",
filename="image_classifier.pt"
)
print(model_path)
You can also download the entire repository:
from huggingface_hub import snapshot_download
model_dir = snapshot_download(
repo_id="hur03/capturemate-image-classifier-v1-5class"
)
print(model_dir)
This model uses a custom PyTorch classifier class from the CaptureMate AI implementation.
It is not a standard AutoModelForImageClassification checkpoint.
The model architecture must match the architecture used during training before loading image_classifier.pt.
The general loading process is:
import torch
from huggingface_hub import hf_hub_download
model_path = hf_hub_download(
repo_id="hur03/capturemate-image-classifier-v1-5class",
filename="image_classifier.pt"
)
# Import the custom classifier implementation
# from your CaptureMate AI codebase.
from your_module import CLIPUIClassifier
model = CLIPUIClassifier(num_labels=5)
state_dict = torch.load(
model_path,
map_location="cpu"
)
model.load_state_dict(state_dict)
model.eval()
Replace:
from your_module import CLIPUIClassifier
with the actual import path of the CLIPUIClassifier implementation from the CaptureMate AI codebase.
The model is based on:
openai/clip-vit-base-patch32
Images should therefore be processed using the corresponding CLIP image processor.
from transformers import CLIPImageProcessor
processor = CLIPImageProcessor.from_pretrained(
"openai/clip-vit-base-patch32"
)
Example:
from PIL import Image
from transformers import CLIPImageProcessor
processor = CLIPImageProcessor.from_pretrained(
"openai/clip-vit-base-patch32"
)
image = Image.open("example_screenshot.jpg").convert("RGB")
inputs = processor(
images=image,
return_tensors="pt"
)
pixel_values = inputs["pixel_values"]
Once the custom CLIPUIClassifier architecture is available and the checkpoint has been loaded, inference can be performed with PyTorch.
import torch
from PIL import Image
from transformers import CLIPImageProcessor
from huggingface_hub import hf_hub_download
from your_module import CLIPUIClassifier
MODEL_ID = "hur03/capturemate-image-classifier-v1-5class"
LABELS = [
"schedule",
"shopping",
"place",
"memo",
"unknown"
]
# Download checkpoint
model_path = hf_hub_download(
repo_id=MODEL_ID,
filename="image_classifier.pt"
)
# Load image processor
processor = CLIPImageProcessor.from_pretrained(
"openai/clip-vit-base-patch32"
)
# Initialize model
model = CLIPUIClassifier(num_labels=5)
state_dict = torch.load(
model_path,
map_location="cpu"
)
model.load_state_dict(state_dict)
model.eval()
# Load screenshot
image = Image.open("example_screenshot.jpg").convert("RGB")
inputs = processor(
images=image,
return_tensors="pt"
)
# Inference
with torch.no_grad():
outputs = model(inputs["pixel_values"])
# Adjust this line if the custom model returns
# an object containing logits instead of raw logits.
logits = outputs
probs = torch.softmax(logits, dim=-1)
pred_id = probs.argmax(dim=-1).item()
confidence = probs[0][pred_id].item()
result = {
"label": LABELS[pred_id],
"confidence": confidence
}
print(result)
Example output format:
{
"label": "shopping",
"confidence": 0.XX
}
The exact inference code may depend on the implementation of
CLIPUIClassifier. The custom model architecture must match the architecture used when the checkpoint was trained.
The model was trained on screenshot images collected for the CaptureMate project.
The same screenshot category structure used by the OCR text classifier was used for image classification.
Dataset split:
| Split | Samples |
|---|---|
| Train | 415 |
| Validation | 89 |
| Test | 90 |
The five target categories are:
| Label | Description |
|---|---|
| schedule | Schedule, reservation, ticket, event, or date-related screenshots |
| shopping | Shopping, product, price, payment, or commerce-related screenshots |
| place | Place, map, restaurant, store, travel, or location-related screenshots |
| memo | Text, article, note, content, or general information screenshots |
| unknown | Ambiguous or non-actionable screenshots |
The image classifier was first evaluated independently on the CaptureMate test set.
| Metric | Value |
|---|---|
| Accuracy | 80.00% |
| Macro F1 | 76.34% |
The image-only classifier performs below the OCR text classifier, which supports its intended role as a complementary visual signal rather than the primary classifier.
The image classifier was also evaluated as part of the CaptureMate multimodal classification pipeline.
The current fixed fusion uses:
Text weight = 0.7
Image weight = 0.3
The resulting multimodal performance is:
| Metric | Value |
|---|---|
| Accuracy | 90.00% |
| Macro F1 | 89.19% |
| Macro Precision | 89.63% |
| Macro Recall | 89.15% |
The multimodal result combines predictions from both the OCR text classifier and this image classifier.
Therefore, the multimodal metrics above should not be interpreted as the standalone performance of this image model.
The model is trained on a relatively small CaptureMate-specific screenshot dataset and has several important limitations.
Known limitations:
unknown, shopping, and memo cases may be confused.This model is recommended as a complementary signal within a multimodal screenshot classification system.
For CaptureMate, the preferred usage is:
OCR Text
↓
Text Classifier
↓
Text Scores ─────────────┐
│
v
Score Fusion
^
│
Image Scores ────────────┘
↑
Image Classifier
↑
Screenshot
The current CaptureMate implementation gives more weight to the OCR text classifier:
Text = 70%
Image = 30%
This reflects the stronger standalone performance of the OCR text model on the current dataset.
The fusion weights may be adjusted in future versions based on validation experiments and additional training data.
This repository contains the files required to distribute the trained image classifier.
image_classifier.ptPyTorch checkpoint containing the trained CLIP vision encoder and custom classifier head weights.
preprocessor_config.jsonConfiguration for image preprocessing based on the CLIP image processor.
Because the model uses a custom PyTorch architecture, the checkpoint alone does not define the complete Python model class.
The corresponding CLIPUIClassifier implementation is required to reconstruct the architecture before loading the weights.
Screenshot Image
|
v
CLIP Image Processor
|
v
CLIP Vision Encoder
(openai/clip-vit-base-patch32)
|
v
Visual Representation
|
v
Custom Classification Head
|
v
5-Class Logits
|
v
Softmax
|
v
Category Prediction
0 -> schedule
1 -> shopping
2 -> place
3 -> memo
4 -> unknown
CaptureMate uses two complementary classification models.
hur03/capturemate-category-classifier-v1-5class
The text classifier analyzes OCR-extracted screenshot text using a fine-tuned KLUE-RoBERTa model.
hur03/capturemate-image-classifier-v1-5class
This repository contains the image classifier, which analyzes screenshot visual information using a CLIP-based architecture.
Together, the two models provide complementary information:
KLUE-RoBERTa
OCR/Text understanding
+
CLIP
Visual understanding
↓
Multimodal screenshot classification
The current CaptureMate fusion strategy combines the two model outputs using fixed weighted score fusion.
The model was developed using:
CaptureMate