Downloads · 30 days
0
Aryan006/cone-distance
cone-distance is a machine learning model from Aryan006. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Detect cones, barriers and stop signs from a single RGB camera, and estimate the distance to each from focal length and pixel geometry. No depth sensor and no learned depth model at inference time — LiDAR appears only…
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2026
Repo size
57.6 MB
Likes
0
Public
Click a slice to open those files.
.pt21.5 MB · 35%
From the Hugging Face model README
Detect cones, barriers and stop signs from a single RGB camera, and estimate the distance to each from focal length and pixel geometry. No depth sensor and no learned depth model at inference time — LiDAR appears only as scoring ground truth, never in the inference path.
This directory is phase 1: get every source into one schema with calibration and ground-truth distance attached, verify the geometry by eye and by assert, and train the detector.
uv venv && source .venv/bin/activate
uv pip install -e . # manifest, merge, export, QA
uv pip install -e ".[nuscenes]" # + the nuScenes extractor
uv pip install -e ".[hub]" # + Hugging Face Hub push/pull
uv pip install -e ".[train]" # + Ultralytics
The AV2, COCO and BDD extractors need no extra dependencies.
Everything runs as a module from the repo root, e.g. python -m src.data.merge.
The sample viewer goes immediately after the first extractor, before writing or running anything else. A sign error or an axis-order mistake in the 3D→2D projection produces plausible-looking garbage that stays invisible until the depth numbers make no sense a day later.
# 0. fetch nuScenes. Presigned URLs expire in minutes -- copy from the browser
# and run immediately. See "Getting nuScenes" below.
./scripts/fetch_nuscenes.sh "$URL" data/raw/nuscenes 'v1.0-mini/*' 'samples/CAM_FRONT/*'
# 1. nuScenes: 2D boxes projected from the 3D cuboids, plus gt distance and calibration
python -m src.data.extract_nuscenes --dataroot data/raw/nuscenes --version v1.0-mini
# 2. look at it, now
python -m src.data.view_samples \
--manifest data/unified/parts/nuscenes.parquet --only-annotated --out qa/nusc.png
# 3. the other sources (any order; AV2 is the droppable one)
python -m src.data.extract_av2 --dataroot /data/av2/sensor/train --max-logs 20
python -m src.data.extract_coco --dataroot /data/coco --split train2017
python -m src.data.extract_bdd_negatives \
--dataroot /data/bdd100k --labels /data/bdd100k/labels/det_20/det_train.json --count 3000
# 4. merge: filter, split by scene, assert the geometry agrees across datasets
python -m src.data.merge
# 5. export
python -m src.data.to_yolo
# 6. publish, then train from the published copy
python -m src.hub.push --repo-id <user>/cone-distance-v1 --private
python -m src.hub.pull --repo-id <user>/cone-distance-v1 --dest data/hub
python -m src.train.train --data data/hub/data.yaml --models yolo11n.pt yolo11s.pt
UNIFIED_ROOT overrides data/unified for every step, or pass --unified-root.
To check the plumbing before any data has landed:
python -m tests.smoke_test
It builds a synthetic ground plane, projects cones onto it, and runs the real merge → export → QA chain over the result.
data/unified/
manifest.parquet # the source of truth: one row per object
parts/<source>.parquet # per-extractor fragments, concatenated by merge.py
calib/<sensor_id>.yaml # K, camera height, ego<-cam rotation
class_priors.yaml # per-class physical size, derived from the manifest
images/<source> # one symlink per dataset, never a copy
yolo/ # derived export: symlinked images + .txt + data.yaml
src/common/ schema.py geometry.py calib.py paths.py
src/data/ extract_*.py merge.py to_yolo.py view_samples.py check_ground_plane.py
src/bench/ pipeline.py eval_pt.py eval_onnx.py eval_ncnn.py compare.py
src/depth/ ground_plane.py known_size.py estimate.py evaluate.py
accuracy.py view_detections.py
src/hub/ push.py pull.py
src/train/ train.py wandb_logger.py checkpoint_mirror.py
scripts/ fetch_nuscenes.sh fetch_coco_stopsigns.py fetch_av2.py make_evalpack.py make_evalpack.py
notebooks/ colab_train.ipynb local_train.ipynb
nuScenes has no download API. Log in at https://www.nuscenes.org/nuscenes#download, right-click a file and copy the link address — it is a presigned CloudFront URL that expires within minutes.
Download with scripts/fetch_nuscenes.sh, which
pipes curl straight into tar and keeps only the files we use. Each trainval blob
is ~30 GB of six cameras, LiDAR and five radars; we use one camera. Unpacking one
to disk needs ~60 GB of headroom for ~450 MB of useful data, so filtering in the
stream is worth the one extra script.
# mini first -- build and verify the pipeline on it
./scripts/fetch_nuscenes.sh "$URL" data/raw/nuscenes 'v1.0-mini/*' 'samples/CAM_FRONT/*'
# then annotations for all 850 scenes, and one blob of images
./scripts/fetch_nuscenes.sh "$URL" data/raw/nuscenes 'v1.0-trainval/*'
./scripts/fetch_nuscenes.sh "$URL" data/raw/nuscenes 'samples/CAM_FRONT/*'
The metadata covers all 850 scenes while each blob holds a slice of the images,
so the extractor skips samples whose blob you did not download and reports the
count as image_not_downloaded.
Extraction, merging and export are CLI scripts and run locally; there is no notebook for them. Training has two, and they train the same thing:
| Notebook | Runs on | Dataset | If it is interrupted |
|---|---|---|---|
colab_train.ipynb | Colab, T4 or better | pulled from the Hub | CheckpointMirror copies last.pt to Drive every 10 epochs, because the runtime's disk does not survive |
local_train.ipynb | any machine — it clones the repo and installs deps if they are not already there, and detects an existing checkout if they are | the local export if present, else pulled | Ultralytics' own last.pt is already on durable disk; section 7 resumes the same run from it |
models, epochs, imgsz and batch are identical in both — yolo11n.pt,
200, 640, 16 — so runs from either are directly comparable. Only the data path,
the device and the run directory differ.
Both the code and the dataset are published to the Hugging Face Hub, because a Colab runtime is ephemeral and re-downloading 18 GB of nuScenes every session is not worth it when the derived dataset is 2.4 GB:
| Repo | What |
|---|---|
Aryan006/cone-distance | this code, and trained weights under weights/ |
Aryan006/cone-distance-v1 | the dataset: shards, manifest, calibration, priors |
Both are private. The Colab notebook reads HF_TOKEN and WANDB_API_KEY from
Colab Secrets; the local one reads them from the environment, then from a
gitignored file of that name in the repo root, then by prompting. Neither needs
editing.
src/hub/push.py packs the YOLO export into ~1 GB tar shards — the Hub handles a
few large files far better than 40,000 small ones, and a shard is a resumable
unit if the upload drops — and uploads manifest.parquet, calib/,
class_priors.yaml and a generated dataset card alongside them. The manifest
stays previewable in the Hub viewer. src/hub/pull.py is the exact inverse, and
rewrites data.yaml with an absolute path so training works from anywhere.
Only the local build uses symlinks. push.py dereferences them, so a shard holds
real bytes and not a link into a nuScenes root that will not exist on the next
runtime.
.txt files.txt files cannot hold intrinsics or ground-truth distance, and phase 2 needs
both. The manifest is the artifact; to_yolo.py is a pure function of it, so the
YOLO directory can be deleted and regenerated at any time.
2D boxes are generated from the 3D cuboids, not taken from any pre-made 2D
annotation file. That projection has to happen anyway to get gt_distance_m, so
doing it for the training labels too means the boxes and the distances come from
one convention, in the same domain the depth estimator is evaluated in.
gt_distance_m is forward depth along the camera optical axis — the Z
component of the object centre in the camera frame — not Euclidean range. Phase 2
must compare like with like.
Barrel is not barrier. An AV2 CONSTRUCTION_BARREL is an orange drum; a
nuScenes barrier is a jersey/fence barrier. Visually unrelated. They stay
separate classes; merging them would tank the class and look like a training
problem rather than a labelling problem.
Split by scene, per source. Adjacent video frames are near-duplicates, so a frame-level random split leaks the val set into training and inflates val mAP a long way. Splitting scenes independently within each source also guarantees every source reaches val — otherwise a random draw can leave val with no nuScenes cones, which is the one thing phase 2 needs.
BDD100K negatives exclude frames containing a traffic sign. BDD's ten classes
have no cones and no barriers, so its frames are safe negatives for those — but
its generic traffic sign bucket hides real stop signs, and pairing an image
containing a stop sign with an empty label file trains the model to suppress stop
signs. --labels enables that filter; without it the extractor warns.
Images that lose every object to the filters are kept as negatives, not
dropped. Everything the filters remove is below the size we intend to detect at
all, so training the model not to fire there matches the deployment target.
--drop-filtered-frames picks the other behaviour.
All in src/common/schema.py, applied in merge.py:
| Threshold | Value | Why |
|---|---|---|
| minimum box height | 15 px | a 6-pixel cone is an unlearnable label that only adds false-positive pressure |
| minimum visibility | 0.4 | drops the nuScenes v0-40 bucket only |
| maximum truncation | 0.5 | above this the box is mostly a guess about what is off-screen |
merge.py assertsThese are real asserts, not things to eyeball. A coordinate-convention mismatch between nuScenes and AV2 produces a manifest that trains YOLO perfectly well and makes the depth numbers nonsense a day later.
gt_distance_m is always positiveThe projected 3D centre landing inside its own 2D box is asserted in the extractors, where the 3D box is still in hand.
Both were found by scoring the ground-plane estimator against gt_distance_m
before anything was built on it. Both would have produced a manifest that trains
YOLO perfectly well and silently wrong distances a day later.
AV2 ships k1, k2, k3 alongside the intrinsics — around -0.24, -0.22, +0.34
for ring_front_center — and the images are not rectified. Projecting through
K alone puts boxes tens of pixels off the objects they label. nuScenes ships
rectified images and passes zeros, for which the model is the identity, so the
same code path now serves both. Calibration.project() and Calibration.ray()
own the forward and inverse; nothing else touches intrinsics directly.
nuScenes puts its ego origin on the ground, so the camera height is simply
the sensor's z translation. Argoverse 2 puts its origin at the rear axle
centre, about 0.26 m up. Taking tz_m as camera height for both makes every
AV2 ground-plane distance read ~18% short, at every range:
| Range | MAE before | MAE after | Bias before | Bias after |
|---|---|---|---|---|
| 0–10 m | 1.48 m | 0.56 m | −1.50 m | +0.11 m |
| 10–20 m | 2.77 m | 1.05 m | −2.49 m | +0.04 m |
| 20–30 m | 5.25 m | 2.39 m | −4.65 m | +0.14 m |
| 30–45 m | 8.51 m | 6.55 m | −7.71 m | +0.38 m |
| 45 m+ | 20.27 m | 18.23 m | −12.52 m | +1.30 m |
Rather than hardcode a vehicle constant, extract_av2.py measures it: cones and
barrels rest on the road, so the median underside of their cuboids in the near
field is the road surface. Applied per log, that also absorbs road grade. The
same derivation is deliberately not applied to nuScenes — there it returns
1.412 m against a stated 1.511 m that the ray-plane fit independently confirms
(implied/stated = 0.986), so the stated value is the better one.
AV2 also gets one sensor_id per log, because its calibration varies between
logs: across these fifteen, focal length moves 1% and camera height 6%. A single
shared av2_ring_front_center would hand phase 2 one arbitrary log's numbers for
every image.
AV2 annotates a stop sign as one cuboid from the ground to the top of the sign —
median 3.22 m tall by 0.86 m wide — so projecting it whole gives a 2D box with a
median aspect ratio of 3.95. COCO's boxes are tight on the octagon at
1.07. Training on both teaches the detector two incompatible definitions of
one class, and a pole-inclusive box is the wrong input to a width-based distance
estimator. A stop sign is a regular octagon, so extract_av2.py takes the
cuboid's own width as the face height and keeps the top slice — no assumed sign
standard. Aspect ratio after: 0.98, and the derived prior drops from 3.21 m
to 0.90 m.
nuScenes v1.0-trainval (metadata for all 850 scenes, images from blob 01),
Argoverse 2 (15 logs), COCO stop signs. 10,261 frames, 23,829 objects.
nuScenes cam_height_m reads 1.511 m.
| Source | barrel | barrier | cone | stop_sign |
|---|---|---|---|---|
| nuscenes | — | 3,224 | 3,112 | — |
| av2 | 11,295 | — | 3,050 | 1,362 |
| coco | — | — | — | 1,786 |
Val split: 287 barrels, 494 barriers, 284 cones, 60 stop signs with
gt_distance_m.
Across all 850 trainval scenes, unfiltered, straight from
sample_annotation.json:
| Class | Height (mean) | Median | 5–95% | n |
|---|---|---|---|---|
| cone | 1.07 m | 1.11 m | 0.70–1.42 | 97,959 |
| barrier | 0.98 m | 0.98 m | 0.78–1.22 | 152,087 |
Worth recording how this went, because the trap is a general one. Measured on
v1.0-mini alone, cones came out at 0.78 m — and mini's ten scenes simply
contain unusually short cones. A prior fitted there would have been 27% low, and
it looked authoritative because it was derived from data rather than looked up.
Deriving beats guessing only once the sample is representative; n_samples is in
class_priors.yaml for exactly this reason, and it is worth reading before
trusting the number above it.
Note also that the derived prior from our filtered manifest (cone 0.97 m) sits below the population figure, because the filters keep the closer, larger, less occluded instances. For phase 2 that is arguably the right bias — the estimator only ever sees objects the detector found, which are drawn from the same skewed distribution — but it is a choice, not an accident.
Recovering distance from the box bottom row using only calib/ — phase 2's
estimator run early, scored against gt_distance_m. Gate: visibility >= 0.7 and
box bottom >= 20 px below the horizon row.
| Range | nuScenes MAE | nuScenes bias | AV2 MAE | AV2 bias |
|---|---|---|---|---|
| 0–10 m | 0.76 m | −0.42 m | 0.56 m | +0.11 m |
| 10–20 m | 2.62 m | −0.58 m | 1.05 m | +0.04 m |
| 20–30 m | 5.87 m | −0.54 m | 2.39 m | +0.14 m |
| 30–45 m | 10.33 m | −4.26 m | 6.55 m | +0.38 m |
| 45 m+ | 17.67 m | −16.18 m | 18.23 m | +1.30 m |
n = 5,222 (nuScenes) and 12,271 (AV2).
Two things this settles. The geometry is correct: sub-metre error in the near field and near-zero bias out to 30 m on both datasets, derived independently from two different calibration sources, is not something a broken transform produces. Box bottoms sit at the ground contact point — measured directly against the cuboids' own bottom faces, the 2D box bottom lands within 1.9 px of the true contact row — so there is nothing to fix in the extractor.
The far field is a conditioning problem, not an accuracy problem.
Sensitivity is Z^2/(fy*h): 0.03 m/px at 10 m, 1.8 m/px at 45 m. The worst
estimates come from boxes bottoming out within a few pixels of the horizon row,
where the ground ray runs nearly parallel to the plane and the intersection
diverges — before the validity gate, single estimates of +3113 m and -2831 m.
Phase 2 therefore needs a horizon-margin guard before it needs a better
estimator, and the known-size cross-check earns its place above roughly 30 m.
AV2 degrades more gracefully at range than nuScenes because its camera has
higher angular resolution (fy 1778 vs 1266) in a portrait frame.
src/train/train.py is plumbing around YOLO.train(), not a modified training
loop. Verified against the reference YOLO11 fine-tuning notebook, which trains
via the CLI:
yolo task=detect mode=train model=yolo11s.pt data=.../data.yaml epochs=10 imgsz=640 plots=True
Resolving both through ultralytics.cfg.get_cfg on the pinned 8.3.40 gives 105
config keys, of which four differ, all plumbing:
| Key | Reference | Here | What it controls |
|---|---|---|---|
device | None | 0 | which GPU |
project | None | /content/runs | where runs are written |
name | None | yolo11n / yolo11s | run directory name |
exist_ok | False | True | overwrite an existing run dir |
All 55 recipe keys are identical: optimiser, lr0, lrf, momentum, weight
decay, warmup, the box/cls/dfl loss gains, nbs, cos_lr,
close_mosaic, patience, and every augmentation (hsv_*, degrees,
translate, scale, shear, perspective, flipud, fliplr, mosaic,
mixup, copy_paste, auto_augment, erasing). Validation matches too —
only device differs, and Model.train() reloads best.pt on completion
([model.py:808]), so model.val() scores the same checkpoint the reference
validates.
This one is not cosmetic. Ultralytics normalises the optimiser step to a nominal batch of 64 by accumulating gradients, and rescales weight decay to compensate:
accumulate = max(round(nbs / batch), 1)
weight_decay = weight_decay * batch * accumulate / nbs
The compensation is exact only when the batch divides 64. batch=-1 hands the
choice to AutoBatch, which computes b = int((f * fraction - p[1]) / p[0]) from
free GPU memory — an arbitrary integer:
| batch | accumulate | effective weight decay |
|---|---|---|
| 16 / 32 / 64 | 4 / 2 / 1 | 1.000× |
| 37 | 2 | 1.156× (+15.6%) |
| 48 | 1 | 0.750× (−25.0%) |
| 100 | 1 | 1.562× (+56.2%) |
So --batch -1 would have quietly trained a different recipe depending on which
GPU Colab handed out, and the run would not be comparable to the reference or to
itself across sessions. The default is now 16 — Ultralytics' own default and the
reference's — and train.py warns if a batch is passed that does not divide 64.
Ultralytics is pinned to <=8.3.40, the version this comparison was run against
and the one the reference notebook pins, so a future default change cannot
silently invalidate the table above.
--wandb logs per epoch, all on the same step so the curves overlay:
| Series | Source |
|---|---|
train/box_loss, train/cls_loss, train/dfl_loss | trainer.tloss |
val/box_loss, val/cls_loss, val/dfl_loss | trainer.metrics |
metrics/* | precision, recall, mAP50, mAP50-95 |
lr/pg0..2 | per param group |
grad_norm/mean, /max, /p95, /clipped_fraction | see below |
Authenticate with $WANDB_API_KEY. Each model gets its own run, so the YOLO11n
and YOLO11s arms are two comparable runs rather than one interleaved mess.
Ultralytics ships its own W&B integration, but it is gated behind
SETTINGS["wandb"], which defaults to False and is persisted to the user's
global config — enabling it reconfigures their machine, not just the run. It
also does not log gradient norm. Since a module was needed for that anyway,
src/train/wandb_logger.py registers its own
callbacks: one file states exactly which series are logged, and nothing outside
the process is changed.
This is the one part of the logging that is not a callback, and it is opt-in
separately as --log-grad-norm. Losses and metrics come from callbacks that only
read trainer.*; without this flag nothing in torch is touched.
BaseTrainer.optimizer_step calls clip_grad_norm_, whose return value is the
total pre-clip gradient norm, and discards it. It then calls zero_grad() in the
same method, and no callback fires between the two — the nearest,
on_train_batch_end, runs after the gradients are already gone, so it cannot be
recomputed either. Wrapping clip_grad_norm_ for the duration of training is the
only hook available, and the most accurate one: the value is captured after
scaler.unscale_(), so it is a true unscaled norm and the same number the
optimiser acted on.
The wrapper only reads a return value. It never touches gradients, the optimiser or the loss, so it cannot change what the model learns. Three things it could still have cost you, all handled:
| Risk | Handling |
|---|---|
on_train_end does not fire when training raises, so the patch survives and the second model's probe wraps the wrapper | install() marks the function it creates and refuses to wrap a marked one; train.py calls close() from a finally |
float() on a CUDA tensor synchronises the device, and clip_grad_norm_ does not otherwise sync (error_if_nonfinite is False), so converting per step would stall every step | norms are kept as detached tensors and converted once per epoch |
Coupling it to --wandb would force the patch on anyone wanting loss curves | separate flag |
tests/test_grad_norm_probe.py pins all of it
against a stub torch — no GPU, no 2 GB install — covering restoration, refusal to
nest, pass-through of the caller's value, and that constructing a probe without
installing it leaves torch untouched.
grad_norm/clipped_fraction is the series to read first: the share of steps
whose norm exceeded the clip of 10. Near 1.0 during warmup is normal. If it stays
there after warmup, the effective step size is set by the clip rather than by
lr0, and lowering the learning rate will do more than tuning anything else.
Verified on a real 3-epoch CPU run against a slice of the actual dataset, with
W&B in offline mode so nothing reached the account. All twelve series arrived;
gradient norm read mean 419 → 1437 → 1423 across the three epochs at
clipped_fraction 1.0, which is warmup behaviour on 48 images.
Colab kills runtimes without warning, and Ultralytics rewrites
weights/last.pt in place every epoch — so a run that dies at epoch 140 leaves
one mutable file and no history.
src/train/checkpoint_mirror.py copies
last.pt out to Drive every N epochs under an epoch-stamped name.
It polls rather than hooking into training, so a failure in it cannot take the
run down. It reads the epoch counter out of the checkpoint rather than counting
its own ticks, so it stays correct across a restart and ignores the epoch: -1
that strip_optimizer writes when training finishes. And it verifies the
copy, not the source, before os.replace-ing it into place: Ultralytics can
begin rewriting last.pt between the epoch check and the copy, so the source
having been loadable a moment ago says nothing about what actually landed. A
name in the destination therefore only ever appears once the file behind it is
whole.
It lives in the repo rather than in a notebook cell so that
tests/test_checkpoint_mirror.py can cover
it against a stub torch — no GPU, no 2 GB install — including the torn-copy race,
the duplicate-epoch guard, and the stripped final checkpoint.
src/depth/ground_plane.py ray-plane intersection. THE REFERENCE TO TRANSCRIBE.
src/depth/known_size.py distance from apparent size, given a size prior
src/depth/estimate.py box -> distance: sample, robust summary, class routing
src/depth/evaluate.py score against LiDAR ground truth on nuScenes val
ground_plane.ground_depth is written to be read and rewritten — the maths is
laid out a step at a time, with the two traps named (h is the height above the
road, not a translation component; Argoverse 2 pixels need undistorting first).
759 objects, 100% given an estimate.
| Range | n | MAE | bias | p90 abs | rel |
|---|---|---|---|---|---|
| 0–10 m | 103 | 1.05 m | −1.22 m | 1.48 m | 14.1% |
| 10–20 m | 339 | 3.96 m | −2.14 m | 7.65 m | 22.9% |
| 20–30 m | 147 | 4.42 m | +0.72 m | 8.64 m | 17.9% |
| 30–45 m | 124 | 8.97 m | −4.61 m | 16.20 m | 24.4% |
| 45 m+ | 46 | 12.13 m | −9.77 m | 20.66 m | 22.7% |
Reading the per-class row first is misleading: cones come out at −5.91 m bias and barriers at −1.23 m, which looks like a class problem. It is not. Broken down by scene, barriers and cones inside the same scene agree closely (−0.401 vs −0.351 in one, −0.059 vs −0.061 in another). The apparent cone bias is an artefact of one scene holding 120 of the 265 cones.
Median relative error by scene runs from −35% to +35%, standard deviation 24%. Within a scene it is close to a constant fraction, which is the signature of the flat-road assumption failing on a grade rather than of a noisy estimator.
Fitting a single constant row shift per scene — one pitch offset, nothing per-object — halves the error:
| Scene | n | shift | median abs error |
|---|---|---|---|
| c08e31a5 | 160 | −50.0 px | 6.56 → 1.30 m |
| 638bfb0a | 97 | −36.9 px | 1.53 → 0.42 m |
| cdb711af | 52 | −5.5 px | 4.50 → 3.23 m |
| fc01a4d7 | 295 | −0.8 px | 1.98 → 1.91 m |
| 79f1ca3a | 75 | +4.1 px | 2.10 → 1.70 m |
| mean | 3.34 → 1.71 m |
−50 px at fy = 1266 is an effective pitch error of about 2.3°, which is an ordinary road grade or a loaded suspension. The ground-plane and known-size per-scene errors correlate only +0.25, so they fail independently — this is the geometry, not a ground-truth artefact.
The conclusion for the rest of phase 2 is that per-frame ground-plane estimation (horizon or vanishing-point) is worth more than any refinement of the per-object maths. The static extrinsic is the binding constraint.
Distance comes from a band of candidate ground-contact rows centred on the bottom edge, not extended upward into the box. The interior of a box is not the ground — a pixel halfway up a cone is half a metre in the air — so an upward band biases every estimate outward: a synthetic cone at 20.0 m reads 21.5 m with a band over the bottom 20% of the box, and 20.0 m with the band centred.
Centring also makes the median exact. Distance is monotonic in row and the median commutes with monotonic maps, so the median of the sampled distances is exactly the distance at the median row. The band costs nothing in accuracy and buys two things: rays that escape over the horizon drop out as NaN instead of poisoning an average, and the spread of the survivors is a confidence signal.
| Class | Estimator |
|---|---|
| cone, barrier, barrel | ground plane, cross-checked against the height prior |
| stop_sign | known size only — pole-mounted, so the box bottom is not a ground contact |
Requiring the two estimators to agree within 40% keeps 87% of objects and moves MAE from 4.97 m to 4.11 m.
Architecturally complete, numerically meaningless, and worth stating plainly.
Stock YOLO11n knows COCO, which has no cone, barrel or barrier class — of the
four classes here only stop_sign exists in COCO, and nuScenes contributes no
stop signs. On nuScenes val it matched 2 of 759 ground-truth boxes at
IoU ≥ 0.5. Relaxing to conf 0.05 and IoU 0.3 reaches 18 of 300, and those are
coincidental overlaps with cars and people rather than detections of our classes:
over 25 val frames the detector fires on car (54), person (24), traffic light
(24), truck (1) and fire hydrant (1).
So --source gt is the number that means something today, and --source detect
becomes meaningful the moment the fine-tuned weights exist. Keeping them separate
is what makes it possible to tell the depth module's error from the detector's.
Three runners over one measured pipeline:
python -m src.bench.eval_pt --threads 4
python -m src.bench.eval_onnx --threads 4
python -m src.bench.eval_ncnn --threads 4 # on the Pi
python -m src.bench.compare
They differ only in which file they hand to Ultralytics. That is deliberate:
three independently written scripts would drift and make the timings
incomparable. Each writes bench/<tag>_timing.json and a CSV of every detection
with its estimated distance.
| Stage | What |
|---|---|
read | JPEG off disk into a numpy array |
preprocess | letterbox and normalise |
inference | the forward pass — the only stage the format changes |
postprocess | NMS and rescaling boxes |
depth | box → distance, the phase 2 estimator |
Images are read here rather than passing predict() a path, so disk I/O lands
in read instead of inflating preprocess. Five warm-up frames are discarded:
the first inference through any backend pays for lazy allocation and cold
caches, and on NCNN it can be an order of magnitude slower than steady state.
p50 and p95 rather than mean, because the tail is real and a mean reports the
tail rather than the typical frame.
| stage (p50 ms) | PyTorch | ONNX |
|---|---|---|
| read | 3.39 | 3.14 |
| preprocess | 1.33 | 1.78 |
| inference | 54.65 | 32.36 |
| postprocess | 0.52 | 1.03 |
| depth | 1.60 | 1.59 |
| total | 63.20 | 40.22 |
| FPS | 15.82 | 24.86 |
| size | 5.6 MB | 10.6 MB |
ONNX is 1.69× on inference and 1.57× end to end. Note inference is 86% of the PyTorch frame but only 80% of the ONNX one — the faster the model, the more the rest of the pipeline matters, which is the whole reason for splitting the stages.
depth is the control rowIt is identical numpy over the same boxes whatever the backend, so it must be constant across columns. It came out at 1.60 and 1.59 ms, and that agreement is what makes the rest of the table trustworthy.
It did not start that way. Unpinned, it read 1.5 ms under PyTorch and 16.2 ms
under ONNX — a 10× swing in code that never changed. onnxruntime defaults to
intra_op_num_threads = 0, meaning every core, while torch was pinned to
--threads; worse, its pool spin-waits between inferences rather than sleeping,
so it starved the single-threaded depth stage even after CPU affinity capped the
process. Ultralytics builds its session with no SessionOptions, so
pipeline.configure_onnxruntime wraps the constructor to inject
intra_op_num_threads and session.intra_op.allow_spinning = 0.
Without that, the benchmark measures thread budget rather than export format.
compare.py warns when the depth row varies by more than 25% between columns,
because that is the signature of the same mistake returning.
--threads sets CPU affinity for the whole process, not just torch, since no
single library-level setting binds them all. It also makes an x86 run a better
proxy for a Pi 5, which has four cores.
It loads — both files parse, the graph builds, the input is accepted — and then segfaults in the forward pass. Every op in the export is standard, and a trivial two-layer graph runs fine in the same wheel, so this is a version mismatch between the pnnx that wrote the export and the installed ncnn reading it: the weight layout changed between them.
A segfault cannot be caught in-process, so eval_ncnn.py runs one forward pass
in a child process first, where a crash is only a return code, and prints the
diagnosis above instead of taking the benchmark down with no output. Re-export on
the Pi and it should run there; that is the number that matters anyway, since
NCNN's advantage is ARM-specific and largely absent on x86.
0 of the four classes detected, as in phase 2 — COCO has no cone, barrel or
barrier. The detector fires on car, person, truck, traffic light. Every
detection is still routed through the depth estimator so the depth stage is
timed on a realistic detection count (~5.7 per frame), but the distances are
meaningless until the fine-tuned weights land.
python scripts/make_evalpack.py --frames 100 --tar
Builds evalpack/ (40 MB, 37 MB tarred): the three model formats, a fixed set
of frames as real files rather than symlinks, the calibration and priors
those frames need, the depth and benchmark code, and four shell scripts. On the
far end:
./setup.sh # once: venv + pinned deps. Uses uv if present.
TORCH_CPU=1 ./setup.sh # same, without the 2.5 GB CUDA build of torch
./run_all.sh # three formats, then the comparison table
Frames are taken in sorted order rather than sampled, so two builds of the same
size contain the same frames and two machines really are comparable. Only the
calibrations those frames use are copied. The tarball excludes .venv and
results/, since the pack is usually archived after someone has already run
setup in it and a venv built for the wrong architecture is worse than useless
on the far end.
THREADS defaults to 4 — a Pi 5's core count — and should be the same on every
machine being compared, for the reason in the depth row discussion above.
Verified end to end on a clean venv: setup resolved Python 3.11 with torch
2.14, ultralytics 8.3.40, onnxruntime 1.30 and ncnn; run_all.sh produced
PyTorch 34.88 ms and ONNX 21.16 ms inference with the depth control row at
0.98 and 1.01 ms, and NCNN reported its version mismatch without taking the run
down.
Numbers from two machines are only comparable if both ran the same code over the
same frames with the same models. scripts/make_evalpack.py builds a
self-contained package that carries all of it:
python scripts/make_evalpack.py --frames 100 --tar
# ship evalpack.tar.gz, then on the target machine:
./setup.sh # or TORCH_CPU=1 ./setup.sh on a GPU-less x86 box
./run_all.sh
40 MB: three model formats, 100 fixed frames as real files, the calibration and priors those frames need, the depth and benchmark code, and one shell script per format.
Three decisions in the builder worth knowing:
run_all.sh lets NCNN fail without taking the run down, since on x86 it
segfaults on a pnnx/ncnn mismatch and the other two numbers are still worth
having. THREADS defaults to 4, matching a Pi 5, and should be the same on
every machine being compared.
evaluate.py asks whether the depth maths works. accuracy.py asks how good a
particular checkpoint is end to end, and saves the answer so checkpoints can
be compared weeks apart without re-running the earlier one.
python -m src.depth.accuracy --weights trained/yolo11n/weights/best.pt --tag best
Writes results/depth_accuracy/<tag>_detections.csv (one row per detection:
class, confidence, predicted distance, true distance, error) and
<tag>_summary.json (the aggregates plus the metadata needed to know what
produced them). Collect at a low --conf; every threshold in the report is
applied afterwards, so one run covers all operating points.
Detections are matched to ground truth greedily, most confident first, and each ground-truth object can be claimed once. Without that rule two overlapping detections of one cone both count and "objects found" exceeds 100% — which it did before this was fixed. A detection is correct only if it overlaps at IoU ≥ 0.5 and carries the right class.
| keep above | detections | % real | objects found | distance error |
|---|---|---|---|---|
| 0.05 | 2011 | 27% | 72% | 3.7 m |
| 0.25 | 863 | 54% | 61% | 3.5 m |
| 0.50 | 447 | 73% | 43% | 2.1 m |
| 0.60 | 300 | 91% | 36% | 1.6 m |
| 0.75 | 104 | 96% | 13% | 1.1 m |
Confidence predicts distance accuracy, which it has no obvious right to do — it is a statement about the box, not the geometry:
| confidence | n | error | within 10% | within 25% |
|---|---|---|---|---|
| 0.05–0.25 | 83 | 4.2 m | 31% | 80% |
| 0.25–0.50 | 137 | 6.9 m | 21% | 50% |
| 0.50–0.75 | 225 | 5.6 m | 21% | 51% |
| 0.75+ | 100 | 1.1 m | 63% | 98% |
The link is the bottom edge. The model is confident when an object is close and unoccluded, which is exactly when its box bottom sits cleanly on the road — and that pixel is the entire distance estimate. So confidence is a usable proxy for "trust this distance", which is worth knowing when deciding what to act on.
| mAP50 | mAP50-95 | detections at 0.5 | % real | error | |
|---|---|---|---|---|---|
| best (epoch 106) | 0.511 | 0.306 | 447 | 73% | 2.1 m |
| last (epoch 120) | 0.509 | 0.302 | 505 | 69% | 2.5 m |
A 0.004 mAP gap and identical timing — they are the same model. Use best.
Per class, barrel is the outlier: mAP50 0.155 against 0.74 for stop_sign,
despite having 11,295 training instances, more than any other class. AV2 barrels
sit at a median 73 m, so nearly all of them are small boxes near the horizon.
view_detections.py draws detections with predicted and true distance, coloured
by relative error, one frame per scene:
python -m src.depth.view_detections --weights trained/yolo11n/weights/best.pt --only-matched
| Needs | Comes from |
|---|---|
fx, fy, cx, cy | calib/<sensor_id>.yaml |
| camera height, ground normal | same file — Calibration.ground_normal_cam() |
| per-class height/width priors | class_priors.yaml |
| GT distance to score against | manifest.gt_distance_m |
| detections to estimate from | runs/detect/yolo11n/weights/best.pt |
extract_av2.py reads the AV2 feather files directly rather than through the
devkit dataloader. The layout is documented and stable, but it has not yet been
run against real AV2 data — verify the first log's output with
view_samples.py before extracting the rest.sensor_id = coco_unknown and have no calibration file, on
purpose: those images have no shared intrinsics and phase 2 must not try to
estimate distance from them.v1.0-trainval metadata covers all 850 scenes while each
*_blobs.tgz holds only a slice of the images, so downloading one or two blobs
is normal. The extractor skips samples whose image is not on disk and reports
the count as image_not_downloaded.Ultralytics YOLO11 weights are AGPL-3.0.