Downloads · 30 days
0
Sathya77/ViT_MvTec
ViT_MvTec is a image classification model from Sathya77. Use it when you need a label for an image. It is set up for transformers. The card lists the license as apache-2.0.
This model detects defects in metal nuts using Vision Transformer (ViT) embeddings + kNN anomaly scoring. It separates good parts from defective ones by analyzing the distribution of anomaly scores. This project uses…
Downloads · 30 days
0
Access
Public
Updated Sep 30, 2025
Repo size
1.4 MB
Likes
0
Public
Click a slice to open those files.
.ipynb1.8 MB · 57%
From the Hugging Face model README
This model detects defects in metal nuts using Vision Transformer (ViT) embeddings + kNN anomaly scoring.
It separates good parts from defective ones by analyzing the distribution of anomaly scores.
This project uses Vision Transformer (ViT, google/vit-base-patch16-224) to extract robust image embeddings.
Fine-tuned it for anomaly detection with a 4-channel input (RGB + edges).
Embeddings are scored via k-Nearest Neighbors (kNN), where outliers yield higher anomaly scores.
Preprocessing includes 224×224 resizing, Canny edge maps, and normalization.
Evaluation uses anomaly score distribution, thresholding (mean+std), and accuracy metrics.
Visual results highlight Top-K most normal and most anomalous samples for interpretability.
train/ → only good samplestest/ → mixture of good and defective samplesThis model is a Vision Transformer (ViT) + k-Nearest Neighbors (kNN) anomaly detection pipeline trained and evaluated on the MVTec Anomaly Detection dataset.Unlike standard supervised classification, this approach learns feature embeddings of normal and defective objects, then applies distance-based anomaly scoring with kNN. The model is designed to detect subtle manufacturing defects such as:
Scratches, Contamination, Surface defects, Structural anomalies
The core idea is that normal samples cluster closely in embedding space, while defective ones appear as outliers with higher anomaly scores.



from transformers import ViTModel
import joblib, numpy as np, torch
# Load model + embeddings
model = ViTModel.from_pretrained("Sathya77/ViT_MvTec")
knn = joblib.load("knn.pkl")
train_embeddings = np.load("train_embeddings.npy")
# Inference
with torch.no_grad():
outputs = model(pixel_values=batch)
emb = outputs.last_hidden_state[:, 0, :]
dists, _ = knn.kneighbors(emb)
score = dists.mean(axis=1)