Downloads · 30 days
0
robotflowlabs/depth-anything-v2-large
depth-anything-v2-large is a depth estimation model from robotflowlabs. Use it for the depth estimation task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Depth Anything V2 (Large, ViT-L backbone) converted to SafeTensors format for safe, fast loading in robotic depth estimation pipelines. 335M parameters for high-quality monocular depth maps.
Downloads · 30 days
0
Access
Public
Updated Mar 19, 2026
Parameters
335M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.3 GB · 100%
From the Hugging Face model README
Depth Anything V2 (Large, ViT-L backbone) converted to SafeTensors format for safe, fast loading in robotic depth estimation pipelines. 335M parameters for high-quality monocular depth maps.
This model is part of the RobotFlowLabs model library, built for the ANIMA agentic robotics platform — a modular ROS2-native AI system that brings foundation model intelligence to real robots operating in the real world.
Monocular depth estimation is fundamental to robotic navigation and manipulation — robots need to know how far away things are from a single camera. Depth Anything V2 produces the highest-quality relative depth maps from a single image. The original weights are distributed as raw .pth files. We converted them to SafeTensors format for safe, zero-copy memory-mapped loading.
| Property | Value |
|---|---|
| Architecture | DPT head + ViT-Large encoder |
| Parameters | 335M |
| Encoder | ViT-L/14 (DINOv2-based) |
| Input Resolution | Flexible (recommended 518×518) |
| Output | Dense relative depth map |
| Training | Synthetic + real depth labels (multi-stage) |
| Original Model | depth-anything/Depth-Anything-V2-Large |
| License | Apache-2.0 |
depth-anything-v2-large/
├── model.safetensors # 1.3 GB — Full model weights
└── README.md # This file
from safetensors.torch import load_file
import torch
# Load SafeTensors weights
state_dict = load_file("model.safetensors")
# Load into Depth Anything V2 architecture
from depth_anything_v2.dpt import DepthAnythingV2
model = DepthAnythingV2(encoder='vitl', features=256, out_channels=[256, 512, 1024, 1024])
model.load_state_dict(state_dict)
model.to("cuda").eval()
# Predict depth
depth = model.infer_image(image) # Returns relative depth map
from transformers import AutoModelForDepthEstimation, AutoImageProcessor
import torch
processor = AutoImageProcessor.from_pretrained("depth-anything/Depth-Anything-V2-Large")
model = AutoModelForDepthEstimation.from_pretrained("depth-anything/Depth-Anything-V2-Large")
model.to("cuda").eval()
inputs = processor(images=image, return_tensors="pt").to("cuda")
with torch.no_grad():
depth = model(**inputs).predicted_depth
from forge.vision import VisionEncoderRegistry
depth_estimator = VisionEncoderRegistry.load("depth-anything-v2-large")
depth_map = depth_estimator(image_tensor) # Relative depth map
Depth estimation is critical across ANIMA modules:
| Model | Params | Size | Best For |
|---|---|---|---|
| depth-anything-v2-large | 335M | 1.3 GB | Highest quality depth |
| depth-anything-v2-small | 24.8M | 95 MB | Real-time edge deployment |
depth-anything/Depth-Anything-V2-Large by TUM & HKU@article{yang2024depth_anything_v2,
title={Depth Anything V2},
author={Yang, Lihe and Kang, Bingyi and Huang, Zilong and Zhao, Zhen and Xu, Xiaogang and Feng, Jiashi and Zhao, Hengshuang},
journal={arXiv preprint arXiv:2406.09414},
year={2024}
}