Downloads · 30 days
0
Renesas/DINOv2-Base-ONNX
DINOv2-Base-ONNX is a image feature extraction model from Renesas. Use it for the image feature extraction task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Not a classifier. DINOv2 is a self-supervised embedding backbone (ViT-Base). Its output is a lasthiddenstate feature/embedding tensor, not class logits — do not treat this model as an ImageNet classifier.
Downloads · 30 days
0
Access
Public
Updated Oct 9, 2026
Repo size
347 MB
Likes
0
Public
Click a slice to open those files.
.onnx347 MB · 100%
From the Hugging Face model README
Not a classifier. DINOv2 is a self-supervised embedding backbone (ViT-Base). Its output is a
last_hidden_statefeature/embedding tensor, not class logits — do not treat this model as an ImageNet classifier.
This repository hosts DINOv2 (ViT-Base) targeting the Renesas R-Car X5H platform for image feature-extraction (embedding) inference on the NPX6 NPU.
dinov2_baselast_hidden_state embedding tensor, suitable
as input to downstream retrieval, clustering, or fine-tuned classification/segmentation headsThe FP32 ONNX model is auto-cast to INT8 by the Renesas MWMX toolchain at compile time — no separate quantization step is required.
dinov2_base_model_int8_fromHF_fp32.onnx (FP32)
│
└─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU ──▶ last_hidden_state
| Artifact | Status | Notes |
|---|---|---|
| FP32 (ONNX) | ✅ Published | fp32/dinov2_base_model_int8_fromHF_fp32.onnx — auto-cast to INT8 by the MWMX toolchain at compile time (see Deployment Flow above); no separate INT8 file is shipped |
Measured on Renesas R-Car X5H via the MWMX runtime (APM50 ship-performance CI pipeline).
Benchmark configuration: Single NPU · Batch size: 1 · Input resolution: not available from source data — TBD
| Parameters | Runtime | Precision | Device | Latency (ms) | Type |
|---|---|---|---|---|---|
| 86M | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 530.03538 | Measured |
| 86M | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 321.912461 | Measured |
TBD — not applicable in the usual classification-accuracy sense. As an embedding backbone, quality would typically be evaluated via downstream task performance (e.g. k-NN classification, retrieval) or embedding similarity to the FP32 reference — not yet measured/published for this repo.
Measured on R-Car X5H with ONNX Runtime + Renesas Execution Provider (INT8 auto-cast); nodes the NPU cannot run fall back to the CPU EP. Batch size 1, 1 NPU. Source: ORT Renesas EP test report, status 2026-09-30. This section supersedes earlier ORT figures in this card.
| Component | AI Cores | Latency (ms) | Throughput (fps) | NPU Inference (ms) | NPU % | CPU + Overhead % | Portable pkg NPU-only (ms) |
|---|---|---|---|---|---|---|---|
| Full model | 1 | 451.443 | 2.2 | 449.585 | 99.6 % | 0.4 % | 528.402 |
| Full model | 12 | 271.338 | 3.7 | 268.613 | 99.0 % | 1.0 % | 260.76 |
12-core latency is not always lower than 1-core: small models are dominated by CPU-side overhead.
Graph partitioning (NPU vs CPU nodes)
| Component | Total Nodes | NPU Nodes | CPU Nodes (Q/DQ inserted) | NPU % | CPU % |
|---|---|---|---|---|---|
| Full model | 320 | 318 | 2 (2) | 99.4 % | 0.6 % |
Compile notes (from the test workbook)
Subgraph config:
default_settings:
input_resolution: [1, 3, 518, 518]
last_hidden_state — patch/token embedding tensor, not class logitsTo run inference on Renesas R-Car X5H, you need:
hf download Renesas/DINOv2-Base-ONNX --repo-type=model --include "fp32/*"
metawaremx_runtime CI pipeline, "APM50" ship-performance target)