Downloads · 30 days
0
AMD-PAVS-AI/dinov2
dinov2 is a image classification model from AMD-PAVS-AI. Use it when you need a label for an image. It is set up for onnx. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Aug 4, 2026
Repo size
4.4 MB
Likes
0
Public
Click a slice to open those files.
.png4.4 MB · 100%
From the Hugging Face model README

DINOv2 is a self-supervised Vision Transformer evaluated here as an ImageNet linear classifier (ViT-S/B/L/G with a 1000-class head). This repository packages evaluation and inference for image classification using ONNX Runtime, exported and validated for AMD ROCm (via the MIGraphX execution provider) so it runs efficiently on AMD GPUs, CPUs, and NPUs.
This is based on the implementation of DINOv2 found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the dinov2 AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.
Task: Image classification
Dataset: ImageNet validation (~50,000 images, 1,000 classes)
Output metrics: Top-1 accuracy, Top-5 accuracy
Model variants: Default is base (
dinov2_vitb14_lc). Override withMODEL_SIZE=small|base|large|giant(hub names:dinov2_vits14_lc,dinov2_vitb14_lc,dinov2_vitl14_lc,dinov2_vitg14_lc).
This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs. Key points:
| Runtime | Precision | Backend | Hardware | Notes |
|---|---|---|---|---|
| ONNX Runtime | FP32, FP16, BF16, INT8 | CPU Execution Provider | AMD CPU | INT8 available via runtime auto-quantization flags |
| ONNX Runtime | FP32, FP16, BF16, INT8 | MIGraphX Execution Provider (ROCm) | AMD GPU | First run may take longer due to graph compilation |
| ONNX Runtime | FP32, FP16, BF16, INT8 | VitisAI Execution Provider | AMD NPU (Ryzen AI) | INT8 requires AMD Quark calibration (make quantize-npu-int8) |
For setup instructions, evaluation scripts, and custom configuration options, see the dinov2 on GitHub.
Model Type: Self-supervised Vision Transformer (ViT) image classifier with a linear-probe head
Base Model: facebookresearch/dinov2 (DINOv2 ViT-B/14 with linear classification head)
Model Stats:
Higher Top-1 accuracy means a larger fraction of images are classified with the correct label as the top prediction — 100% is perfect, 0% is chance-level for random guessing. Top-5 allows credit when the true class appears anywhere in the model's five highest-scoring labels; values above ~80% Top-1 on ImageNet val are considered strong for this linear-probe setup.
| Metric | Description |
|---|---|
| Top-1 | Primary classification metric — the fraction of images where the highest-scoring class matches ground truth. Strictest single-label score; a wrong top prediction counts as a full miss even if the true class ranked second. |
| Top-5 | Fraction of images where the true class appears in the model's top five predictions. Looser than Top-1 and typically higher; useful when near-miss rankings still indicate the model recognized the object category. |
Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?
📂 View the full project on GitHub
The GitHub repository includes: