Downloads · 30 days
79
100% of all-time downloads
mlx-community/Video-Depth-Anything-Large-MLX
Video-Depth-Anything-Large-MLX is a depth estimation model from mlx-community. Use it for the depth estimation task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as cc-by-nc-4.0.
MLX port of Video Depth Anything (ByteDance, CVPR 2025 highlight): consistent monocular depth estimation for arbitrarily long videos. Converted from the official checkpoint depth-anything/Video-Depth-Anything-Large.
Downloads · 30 days
79
100% of all-time downloads
All-time downloads
79
Public
Parameters
385M
1.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.5 GB · 100%
From the Hugging Face model README
MLX port of Video Depth Anything (ByteDance, CVPR 2025 highlight): consistent monocular depth estimation for arbitrarily long videos. Converted from the official checkpoint depth-anything/Video-Depth-Anything-Large.
Architecture: DINOv2 backbone + DPT head with AnimateDiff-style temporal motion modules. Outputs per-frame depth maps, not text.
from mlx_vlm import load
from mlx_vlm.models.video_depth_anything.generate import (
VideoDepthPredictor,
read_video_frames,
)
model, processor = load("mlx-community/Video-Depth-Anything-Large-MLX")
predictor = VideoDepthPredictor(model, processor)
frames, fps = read_video_frames("input.mp4", max_len=300, target_fps=15)
depths = predictor.infer(frames) # (T, H, W) float32
(B, T, H, W, 3); H and W must be multiples of 14.CC-BY-NC-4.0 (same as the source checkpoint). Non-commercial use only.