Downloads · 30 days
0
iloveass/kompon-image-gate
kompon-image-gate is a machine learning model from iloveass. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as bsd-3-clause.
This is a lightweight, 3-class image classification model paired with a heuristic quality pre-check. It serves as the preliminary "gate" for the Kapuni pipeline, deciding whether an uploaded image is a valid building…
Downloads · 30 days
0
Access
Public
Updated Jul 4, 2026
Repo size
6.2 MB
Likes
0
Public
Click a slice to open those files.
.pth6.2 MB · 100%
From the Hugging Face model README
This is a lightweight, 3-class image classification model paired with a heuristic quality pre-check. It serves as the preliminary "gate" for the Kapuni pipeline, deciding whether an uploaded image is a valid building surface (and thus worth passing to downstream damage-assessment models) or an invalid upload (e.g., road, random objects).
building_surface, road_or_pavement, other)ACCEPT or REJECT).Primary Use Case: To act as a low-latency pre-filter for building damage assessment applications. It ensures that heavy, specialized object detection models (Model A) only process relevant images of walls, columns, beams, or slabs, thereby saving compute and providing immediate, tailored feedback to users who upload irrelevant images.
Out-of-Scope Use Cases:
import torch
import torchvision.transforms as transforms
from torchvision import models
import json
from PIL import Image
# 1. Load config
with open("gate_config.json") as f:
cfg = json.load(f)
# 2. Initialize model architecture
model = models.mobilenet_v3_small(weights=None)
model.classifier[3] = torch.nn.Linear(model.classifier[3].in_features, len(cfg["classes"]))
# 3. Load weights
model.load_state_dict(torch.load("gate_mobilenetv3.pth", map_location="cpu"))
model.eval()
# 4. Prepare image
tf = transforms.Compose([
transforms.Resize((cfg["img_size"], cfg["img_size"])),
transforms.ToTensor(),
transforms.Normalize(cfg["mean"], cfg["std"]),
])
img = Image.open("your_photo.jpg").convert("RGB")
input_tensor = tf(img).unsqueeze(0)
# 5. Predict
with torch.no_grad():
logits = model(input_tensor)
probs = torch.softmax(logits, 1)[0]
pred_idx = int(probs.argmax())
confidence = float(probs[pred_idx])
predicted_class = cfg["classes"][pred_idx]
print(f"Class: {predicted_class}, Confidence: {confidence:.2f}")
The model was fine-tuned on a balanced dataset of approximately 2,500 images per class, sourced from existing public datasets:
building_surface: Sourced from HRCDS. Currently composed primarily of damaged concrete walls and building elements.road_or_pavement: Sourced from RDD2022 (specifically the India/dashcam subset) to serve as a strong hard-negative for crack-like patterns on roads.other: Sourced from COCO (val2017) to capture a wide variety of random objects, scenes, people, and vehicles.Because wrongly rejecting a real building crack is a critical failure for a safety tool, the model implements an "Accept-on-Uncertainty" rule.
TAU (default 0.55), the image is automatically ACCEPTED and passed to the downstream model.REJECTED.building_surface class was trained heavily on damaged concrete (due to the HRCDS source). Clean walls, masonry, or brick might result in lower confidence scores. The Accept-on-Uncertainty rule mitigates this by passing low-confidence images through, but the bias exists natively in the weights.