Downloads · 30 days
0
anonymous-eval/food-recognition
food-recognition is a image classification model from anonymous-eval. Use it when you need a label for an image. It is set up for transformers. The card lists the license as apache-2.0.
DINOv3-food is a food image recognition model fine-tuned from facebook/dinov3-vitl16-pretrain-lvd1689m on TSOTSA-Img, a merged food image dataset built from AFD, FruitVeg-81, Food-101, and UECFood256.
Downloads · 30 days
0
Access
Public
Updated Jul 26, 2026
Parameters
304M
4.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.4 GB · 100%
From the Hugging Face model README
DINOv3-food is a food image recognition model fine-tuned from
facebook/dinov3-vitl16-pretrain-lvd1689m
on TSOTSA-Img, a merged food image dataset built from AFD, FruitVeg-81,
Food-101, and UECFood256.
The model predicts 389 food categories and uses a DINOv3 ViT-L/16 backbone with a lightweight linear classification head.
TSOTSA-Img is the merged dataset used for food recognition in this work. It is split into training and test subsets:
The merged dataset combines food images and labels from:
The selected checkpoint was fine-tuned for 8 epochs.
| Setting | Value |
|---|---|
| Base model | facebook/dinov3-vitl16-pretrain-lvd1689m |
| Backbone | DINOv3 ViT-L/16 |
| Number of labels | 389 |
| Epochs | 8 |
| Batch size | 16 |
| Learning rate | 2e-5 |
| Weight decay | 0.01 |
| Warmup ratio | 0.05 |
| Validation selection | best validation behavior, with emphasis on validation loss |
Validation metrics for the selected run:
| Metric | Value |
|---|---|
| Validation loss | 0.1100 |
| Accuracy | 0.9731 |
| Macro-F1 | 0.9727 |
| Top-5 accuracy | 0.9968 |
Final evaluation was performed on the individual source datasets and on the merged TSOTSA-Img test split.
| Dataset | Accuracy |
|---|---|
| FruitVeg-81 | 0.9976 |
| AFD | 0.9997 |
| Food-101 | 0.9551 |
| UECFood256 | 0.8215 |
| TSOTSA-Img test | 0.9062 |
For the TSOTSA-Img test split:
| Metric | Value |
|---|---|
| Accuracy | 0.9062 |
| Macro-F1 | 0.9072 |
This repository stores a custom backbone-plus-classifier model:
backbone/: DINOv3 backbone saved with transformers.classifier.pt: linear classification head.classifier_config.json: label mappings and classifier metadata.preprocessor_config.json: image preprocessing configuration.Because this model uses a custom wrapper around the DINOv3 backbone, loading it
with AutoModelForImageClassification.from_pretrained(...) is not sufficient.
Use the project loader or reconstruct the wrapper before inference.
Example with the project inference class:
from inference.food_classifier import FoodClassifier
model_dir = "model_saved/finetuning/facebook-dinov3-vitl16-pretrain-lvd1689m/epochs_8"
classifier = FoodClassifier(model_dir)
prediction = classifier.predict("path/to/food_image.jpg")
print(prediction)
Manual loading:
import json
import torch
from transformers import AutoImageProcessor, AutoModel
from finetuning.train_classifier import BackboneImageClassifier
model_dir = "model_saved/finetuning/facebook-dinov3-vitl16-pretrain-lvd1689m/epochs_8"
with open(f"{model_dir}/classifier_config.json", "r", encoding="utf-8") as f:
classifier_config = json.load(f)
id2label = {
int(label_id): label
for label_id, label in classifier_config["id2label"].items()
}
label2id = {
label: int(label_id)
for label, label_id in classifier_config["label2id"].items()
}
backbone = AutoModel.from_pretrained(f"{model_dir}/backbone")
model = BackboneImageClassifier(
backbone=backbone,
num_labels=int(classifier_config["num_labels"]),
id2label=id2label,
label2id=label2id,
)
classifier_state = torch.load(f"{model_dir}/classifier.pt", map_location="cpu")
model.classifier.load_state_dict(classifier_state)
model.eval()
processor = AutoImageProcessor.from_pretrained(model_dir)
This model is intended for food image recognition over the TSOTSA-Img label space. It can be used for research experiments, dataset benchmarking, and food recognition pipelines where the target labels overlap with the 389 supported categories.
classifier_config.json.