Downloads · 30 days
0
galihkjaya/DINOv2-Garbage-Classification
DINOv2-Garbage-Classification is a image classification model from galihkjaya. Use it when you need a label for an image. It is set up for pytorch. The card lists the license as mit.
A fine-tuned DINOv2 ViT-L/14 model for garbage image classification.
Downloads · 30 days
0
Access
Public
Updated Jul 20, 2026
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.pth1.2 GB · 100%
From the Hugging Face model README
A fine-tuned DINOv2 ViT-L/14 model for garbage image classification.
This repository is prepared as a workshop baseline model, allowing participants to quickly start experimenting with transfer learning techniques without training a Vision Transformer from scratch.
The checkpoint can be used directly for inference or as an initialization for further fine-tuning on custom datasets.
This project fine-tunes Meta AI's DINOv2 ViT-L/14 backbone for a three-class garbage classification task.
Unlike conventional CNN-based approaches, DINOv2 learns powerful visual representations through self-supervised learning, making it an excellent backbone for downstream computer vision tasks with limited labeled data.
Participants are encouraged to extend this model with techniques such as:
Input Image (518 × 518)
│
▼
DINOv2 ViT-L/14 Backbone
│
▼
1024-D Feature Vector
│
▼
LayerNorm
│
▼
Linear (1024 → 512)
│
▼
GELU
│
▼
Dropout (0.3)
│
▼
Linear (512 → 3)
│
▼
Prediction
facebookresearch/dinov2dinov2_vitl14Training strategy:
The model classifies garbage into three categories.
| Label | Class |
|---|---|
| 0 | Recyclable |
| 1 | Electronic |
| 2 | Organic |
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() |
| Split | Images |
|---|---|
| Training | 23,873 |
| Validation | 2,653 |
| Test | 1,458 |
| Class | Images |
|---|---|
| Recyclable | 12,567 |
| Electronic | 9,999 |
| Organic | 3,960 |
Training images are augmented using:
Input resolution:
518 × 518
.
├── assets/
│ ├── 0_Recyclable/
│ ├── 1_Electronic/
│ └── 2_Organic/
│
├── best_dinov2.pth
└── README.md
import torch
import torch.nn as nn
from huggingface_hub import hf_hub_download
pth_path = hf_hub_download(
repo_id="galihkjaya/DINOv2-Garbage-Classification",
filename="best_dinov2.pth"
)
class DINOv2Classifier(nn.Module):
def __init__(self, num_classes=3, unfreeze_blocks=4):
super().__init__()
self.backbone = torch.hub.load(
"facebookresearch/dinov2",
"dinov2_vitl14"
)
for p in self.backbone.parameters():
p.requires_grad = False
for block in self.backbone.blocks[-unfreeze_blocks:]:
for p in block.parameters():
p.requires_grad = True
self.head = nn.Sequential(
nn.LayerNorm(1024),
nn.Linear(1024, 512),
nn.GELU(),
nn.Dropout(0.3),
nn.Linear(512, num_classes)
)
def forward(self, x):
features = self.backbone(x)
return self.head(features)
model = DINOv2Classifier()
model.load_state_dict(torch.load(pth_path, map_location="cpu"))
model.eval()
from PIL import Image
from torchvision import transforms
transform = transforms.Compose([
transforms.Resize((518, 518)),
transforms.ToTensor(),
transforms.Normalize(
mean=(0.485,0.456,0.406),
std=(0.229,0.224,0.225)
)
])
labels = {
0: "Recyclable",
1: "Electronic",
2: "Organic"
}
image = Image.open("sample.jpg").convert("RGB")
image = transform(image).unsqueeze(0)
with torch.no_grad():
logits = model(image)
pred = logits.argmax(1).item()
print(labels[pred])
This checkpoint serves as the baseline model. Suggested follow-up experiments include:
The objective is to demonstrate how a strong pretrained visual backbone can be adapted efficiently for downstream classification tasks.
This project builds upon:
@misc{galih2026,
title={DINOv2 Garbage Classification},
author={Galih Kusuma Wijaya},
year={2026},
publisher={Hugging Face},
howpublished={https://huggingface.co/galihkjaya/DINOv2-Garbage-Classification}
}