Downloads · 30 days
0
yuliangguo/depth-any-camera
depth-any-camera is a machine learning model from yuliangguo. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
--- license: cc-by-nc-sa-2.0 --- <div align="center" <h1Depth Any Camera: Zero-Shot Metric Depth Estimation from Any Camera</h1
Downloads · 30 days
0
Access
Public
Updated Mar 13, 2025
Repo size
12.3 GB
Likes
5
Public
Click a slice to open those files.
.pt7.2 GB · 58%
From the Hugging Face model README
Yuliang Guo<sup>1*†</sup> · Sparsh Garg<sup>2†</sup> · S. Mahdi H. Miangoleh<sup>3</sup> · Xinyu Huang<sup>1</sup> · Liu Ren<sup>1</sup>
<sup>1</sup>Bosch Research North America <sup>2</sup>Carnegie Mellon University <sup>3</sup>Simon Fraser University
*corresponding author †equal technical contribution
<a href="https://arxiv.org/abs/2501.02464"><img src='https://img.shields.io/badge/arXiv-Depth Any Camera-red' alt='Paper PDF'></a> <a href='https://yuliangguo.github.io/depth-any-camera/'><img src='https://img.shields.io/badge/Project_Page-Depth Any Camera-green' alt='Project Page'></a> <a href='https://github.com/yuliangguo/depth_any_camera'><img src='https://img.shields.io/badge/Code-Depth Any Camera-blue' alt='Code'></a>
</div> <p align="center"> <img src="teaser.png" alt="teaser" style="width: 70%;"> </p>Depth Any Camera (DAC) is a novel zero-shot metric depth estimation framework that extends a perspective-trained model to handle any type of camera with varying FoVs effectively.
Notably, DAC can be trained exclusively on perspective images, yet it generalizes seamlessly to fisheye and 360 cameras without requiring specialized training data. Key features include:
Tired of collecting new data for specific cameras? DAC maximizes the utility of every existing 3D data for training, regardless of the specific camera types used in new applications.
The zero-shot metric depth estimation results of Depth Any Camera (DAC) are visualized on ScanNet++ fisheye videos and compared to Metric3D-v2. The visualizations of A.Rel error against ground truth highlight the superior performance of DAC.
<p align="center"> <img src="video_scannet++_1.gif" alt="animated" /> </p>Additionally, we showcase DAC's application on 360-degree images, where a single forward pass of depth estimation enables full 3D scene reconstruction.
<p align="center"> <img src="video_matterport3d_1.gif" alt="animated" /> </p>Additional visual results and comparison with the prior SoTA can be found at <a href='https://yuliangguo.github.io/depth-any-camera/'><img src='https://img.shields.io/badge/Project_Page-Depth Any Camera-green' alt='Project Page'></a>
Depth Any Camera performs <b>significantly better</b> than the previous SoTA <b>metric</b> depth estimation models Metric3D-v2 and UniDepth in zero-shot generalization to large FoV camera images given <b>significantly smaller training dataset and model size</b>.
| Method | Training Data Size | Matterport3D (360) | Pano3D-GV2 (360) | ScanNet++ (fisheye) | KITTI360 (fisheye) | ||||
|---|---|---|---|---|---|---|---|---|---|
| AbsRel | $\delta_1$ | AbsRel | $\delta_1$ | AbsRel | $\delta_1$ | AbsRel | $\delta_1$ | ||
| UniDepth-VitL | 3M | 0.7648 | 0.2576 | 0.7892 | 0.2469 | 0.4971 | 0.3638 | 0.2939 | 0.4810 |
| Metric3D-v2-VitL | 16M | 0.2924 | 0.4381 | 0.3070 | 0.4040 | 0.2229 | 0.5360 | 0.1997 | 0.7159 |
| Ours-Resnet101 | 670K-indoor / 130K-outdoor | 0.156 | 0.7727 | 0.1387 | 0.8115 | 0.1323 | 0.8517 | 0.1559 | 0.7858 |
| Ours-SwinL | 670K-indoor / 130K-outdoor | 0.1789 | 0.7231 | 0.1836 | 0.7287 | 0.1282 | 0.8544 | 0.1487 | 0.8222 |
We highlight the best and second best results in bold and italic respectively (better results: AbsRel $\downarrow$ , $\delta_1 \uparrow$).