Downloads · 30 days
0
toqi/camtrap-distance
camtrap-distance is a depth estimation model from toqi. Use it for the depth estimation task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
Trained weights for the paper Reference-Conditioned Distance Intervals on Unseen Camera Traps (Computer Vision for Ecology Workshop, ECCV 2026). The model is a Depth Anything V2 metric-outdoor backbone with a 7-channe…
Downloads · 30 days
0
Access
Public
Updated Aug 27, 2026
Repo size
2.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors2.7 GB · 100%
From the Hugging Face model README
Trained weights for the paper Reference-Conditioned Distance Intervals on Unseen
Camera Traps (Computer Vision for Ecology Workshop, ECCV 2026). The model is a
Depth Anything V2 metric-outdoor backbone with a 7-channel input (target photo,
reference flag photo, reference distance map) and a monotone quantile head. For
each pixel it predicts horizontal ground distance in metres as a median and a
90% interval, [q05, q50, q95].
seed1/best/ config.json, model.safetensors training seed 1
seed2/best/ config.json, model.safetensors training seed 2
seed1/best is the checkpoint to use. seed2/best is the second training run;
the paper reports the mean of the two.
Held-out test cameras of the flag survey (12 of 62 cameras, 801 markers), aligned-reference path.
| checkpoint | MAE (m) | p90 (m) | coverage of the 90% interval |
|---|---|---|---|
| seed 1 | 0.841 | 1.857 | 97.9% |
| seed 2 | 0.813 | 1.777 | 98.1% |
| mean (paper) | 0.827 | 1.817 | 98.0% |
The checkpoint has a 7-channel patch embedding and a wrapped head, so it is
loaded through the code repository rather than AutoModel.
hf download toqi/camtrap-distance --local-dir outputs/paper-ckpt
from src.calibration.ground_plane import load_calibration
from src.network.model import load_checkpoint
from src.network.predict import predict
model = load_checkpoint("outputs/paper-ckpt/seed1/best")
calibration = load_calibration("MAS_CAM04", "IMG_0001")
p = predict(model, "new_photo.jpg", "data/flaglabel-dataset/MAS_CAM04/IMG_0001.JPG", calibration)
p.q05, p.q50, p.q95 # (H, W) arrays, metres
See the code repository for installation, the survey data, and training.
CC BY-NC 4.0. The weights derive from Depth-Anything-V2-Large, which is released under CC BY-NC 4.0. The code is MIT.
@inproceedings{sarker2026reference,
title = {Reference-Conditioned Distance Intervals on Unseen Camera Traps},
author = {Sarker, Toqi Tahamid and Islam, Taminul and Morelock, Seth J.
and Bastille-Rousseau, Guillaume and Ahmed, Khaled R.},
booktitle = {Computer Vision for Ecology Workshop, European Conference on Computer Vision (ECCV)},
year = {2026}
}