Downloads · 30 days
0
blanchon/dinodepth-model
dinodepth-model is a depth estimation model from blanchon. Use it for the depth estimation task on the model card, and read the license before you ship it in a product. It is set up for dinov3-depth-head. The card lists the license as apache-2.0.
Trained Simple Depth Transformer (SDT) decoder heads for zero-shot affine-invariant (relative) monocular depth, reproducing AnyDepth (arXiv:2601.02760) on a frozen DINOv3 backbone. Only the small SDT decoder is traine…
Downloads · 30 days
0
Access
Public
Updated Jun 18, 2026
Repo size
75.6 MB
Likes
0
Public
Click a slice to open those files.
.safetensors75.6 MB · 100%
From the Hugging Face model README
Trained Simple Depth Transformer (SDT) decoder heads for zero-shot affine-invariant (relative) monocular depth, reproducing AnyDepth (arXiv:2601.02760) on a frozen DINOv3 backbone. Only the small SDT decoder is trained; the DINOv3 encoder is frozen and loaded separately from Meta's checkpoints. Two heads are provided in this repo:
| File | Backbone | Decoder params | Train |
|---|---|---|---|
sdt-vitl16.safetensors | DINOv3 ViT-L/16 | 13.4 M | 5 epochs |
sdt-vits16.safetensors | DINOv3 ViT-S/16 | 5.5 M | 10 epochs |
These are decoder weights only (~13/5 M params) — pair each with its matching frozen DINOv3
backbone (facebook/dinov3-vitl16-pretrain-lvd1689m / -vits16-).
AbsRel ↓ / δ1 ↑ on NYUv2 (Eigen 654) and KITTI (Eigen 652), scored with per-image least-squares scale+shift alignment in disparity space, Eigen/Garg crop, 10 m / 80 m cap.
| Model | NYU AbsRel | NYU δ1 | KITTI AbsRel | KITTI δ1 |
|---|---|---|---|---|
| ViT-L/16 + SDT (this repo) | 0.068 | 0.955 | 0.093 | 0.911 |
| ViT-S/16 + SDT (this repo) | 0.091 | 0.917 | 0.115 | 0.852 |
| AnyDepth ViT-L (paper) | 0.060 | — | 0.086 | — |
| AnyDepth ViT-S (paper) | 0.082 | — | 0.102 | — |
A faithful reproduction — ~0.01 AbsRel behind the paper on each benchmark (consistent across both backbones; plausibly the augmentation/data-filtering details AnyDepth underspecifies).
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from dinov3_depth.head import DepthModel, DepthModelConfig
# Frozen DINOv3 ViT-L/16 + (randomly-initialised) SDT head; default config matches the trained head
# (GroupNorm, fusion_channels=256).
model = DepthModel.from_pretrained(DepthModelConfig(backbone="vitl16"))
head = hf_hub_download("blanchon/dinodepth-model", "sdt-vitl16.safetensors")
model.head.load_state_dict(load_file(head))
model.eval()
# images: float [B, 3, H, W] in [0, 1], H and W multiples of 16. Returns affine-invariant disparity.
disparity = model(images)
(Use backbone="vits16" + sdt-vits16.safetensors for the small head.)
blanchon/dinodepth-dataset (Hypersim, VKITTI2,
BlendedMVS, IRS, TartanAir). 768² input, AdamW lr 1e-3, PolyLR.