Downloads · 30 days
0
ucsde4e/fishsense-laser-detector
fishsense-laser-detector is a object detection model from ucsde4e. Use it when you need objects located in an image. It is set up for pytorch. The card lists the license as bsd-3-clause.
Post-processing detector for the laser dot in underwater dive imagery from the FishSense project. Given a 4K underwater dive frame (Olympus TG-6 ORF, 4014 × 3016), returns either (predx, predy, confidence) or no laser…
Downloads · 30 days
0
Access
Public
Updated Jul 24, 2026
Repo size
491 MB
Likes
1
Public
Click a slice to open those files.
.pt491 MB · 100%
From the Hugging Face model README
Post-processing detector for the laser dot in underwater dive imagery from the
FishSense project. Given a 4K
underwater dive frame (Olympus TG-6 ORF, 4014 × 3016), returns either
(pred_x, pred_y, confidence) or no laser detected. The laser dot is used
downstream for depth estimation via laser triangulation to measure fish length.
| file | architecture | val hit_n3 | test hit_n3 | bias offset | notes |
|---|---|---|---|---|---|
run3_epoch_021.pt | ResNet-34 UNet (smp), 6-ch input | 0.9100 | 0.8635 | (−0.179, −0.023) | Production. 24M params. |
run7_hrnet_w18_epoch_021.pt | HRNet-w18 UNet (smp), 6-ch input | 0.9088 | 0.8632 | (−0.436, −0.441) | 16M params. Tied with production; matrix-tested alternative. Bias offset is materially different from run3's — see caveat below. |
Bias offsets are val-inlier signed-residual means under fp32 inference (issue #13). Both refits published as of 2026-07-24.
run7 offset caveat: run7's val-derived bias generalizes to test on hit_n3 (flat within noise) and hit_n4 (+0.56 pp) but degrades median test error from 0.915 px to 1.099 px — the offset over-corrects well-behaved frames while shifting the borderline 3-5 px tail. Per- wavelength on test, run7's residuals show more spread than run3's (green residual sign flips vs. val). For downstream 3D reconstruction where unbiased pixel predictions matter more than median-frame tightness, the val-derived offset is the correct choice. For a downstream that prefers median tightness over unbiased mean, apply no offset to run7. run3 shows neither tension — its val-derived offset transfers cleanly on both metrics.
Both checkpoints are trained on the same data with the same 6-channel input
pipeline (chromaticity + wavelength + Bayer-excess G / R). The two architectures
land within statistical noise of each other on both splits, confirming a
label-noise floor rather than a model-capacity ceiling. See the training
repo's notes/architecture_ablation_matrix.md for the full 10-config ablation.
Numbers below are with the production inference recipe (next section) on the FishSense dive-level splits (26 val dives / 27 test dives).
| split | hit_n3 (within 3 px) | hit_n4 (within 4 px) |
|---|---|---|
| val | 0.9100 | 0.9327 |
| test | 0.8635 | 0.8934 |
Numbers are on Ada RTX 4500 under fp32 inference (issue #13 fix), with
the −0.179 −0.023 bias offset. Previous bf16 numbers were val 0.9081 /
test 0.8615 with −0.200 −0.006; the +0.19–0.20 pp lift comes from
removing the argmax quantization. Fp32 inference is IEEE-754 stable —
you should reproduce these numbers bit-perfect on any hardware, not
just Ada.
These metrics sit at the labeler click-noise floor. Perpendicular
label-to-line σ on clean dives is ≈ 0.77 px (from RANSAC per-dive line fits);
the detector's predictions are tighter than the median human label. The
residual 3-10 px "borderline" failure mode is dominated by click variance,
not model imprecision. See the training repo's notes/imwut_ba_findings.md
for the error-budget analysis.
The checkpoints are pickled dicts with model_state_dict, cfg (a
TrainConfig dataclass), epoch, and per-epoch metrics. Load via the
training repo:
git clone https://github.com/UCSD-E4E/2026-05-02_laser_detector
cd 2026-05-02_laser_detector
# Download the checkpoint
huggingface-cli download ucsde4e/fishsense-laser-detector \
run3_epoch_021.pt \
--local-dir data/phase2/checkpoints_sensor_bayer_50e_run3/
nix develop --command uv run torchrun --standalone --nproc_per_node=4 \
scripts/audit_failures.py \
--checkpoint data/phase2/checkpoints_sensor_bayer_50e_run3/run3_epoch_021.pt \
--split val \
--image-pipeline linear_npy \
--bayer-excess-cache-dir data/image_cache_bayer_excess \
--soft-snap-inference \
--rig-prior --rig-prior-floor 1.0 \
--cascade \
--subpixel-refine \
--line-mask-corridor-px 25 \
--pixel-bias-offset -0.179 -0.023 \
--out-dir data/audit/your_run_name
The production inference recipe applies six inference-side steps:
| flag | purpose |
|---|---|
--soft-snap-inference | Blend prediction toward per-dive line projection (α ≤ 0.3) |
--rig-prior --rig-prior-floor 1.0 | Hard bbox mask on the empirical laser-position distribution in sensor coords |
--cascade | Two-pass inference (coarse tile-grid argmax + 256×256 refined crop) |
--subpixel-refine | Parabolic-fit sub-pixel peak refinement on heatmap logits (not sigmoid probs — see bf16 caveat below) |
--line-mask-corridor-px 25 | Per-dive geometric corridor mask around the fitted dive line |
--pixel-bias-offset -0.179 -0.023 | Constant per-checkpoint calibration; empirically derived from val inliers |
Inference runs fp32 end-to-end. Under torch.autocast(bf16), the
forward pass emits bf16 heatmap logits — and even though we promote to
fp32 before sigmoid, the underlying logit values were already quantized
in bf16 (the .float() cast is lossless zero-padding; precision is
gone). Rival pixels routinely round to identical bf16 buckets, and
flat.max() breaks the tie by row-major index. Because cuDNN picks
different tensor-core kernels on Ada (SM89) vs Ampere (SM80) vs Hopper,
the same weights + same input produce different argmax outcomes on
different GPUs — shifts up to ~200 px on frames with sub-ulp margins.
The training repo's predict_frame and predict_frame_with_cascade
default autocast_dtype=None. Training still uses bf16 for throughput,
but the inference path is fp32 throughout. Any port must match this
default to reproduce the published val 0.9100 / test 0.8635.
If you write new inference code, either wrap the argmax stage in
autocast(dtype=torch.bfloat16, enabled=False) or promote to fp32
before the forward pass. Sub-pixel refinement operates on logits,
not on post-sigmoid probabilities.
Because the current inference path is IEEE-754 fp32 throughout, the published metrics are hardware-independent. A port that matches the recipe should reproduce val 0.9100 / test 0.8635 bit-perfect on any CUDA architecture (Ada, Ampere, Hopper, consumer) or CPU. If your port diverges by more than measurement noise (~0.05 pp on val, ~0.05 pp on test), you likely have bf16 or fp16 autocast leaking into the argmax stage.
By default, predictions are in raw pixel space (same coord frame as the
raw image the detector consumed). Labels in the FishSense corpus are in
rectified (undistorted) pixel space because the labeling UI renders via
RectifiedImage(RawImage(...)). Impact on hit_n3 is negligible (median 0.02
px displacement, p99 1.01 px), but coord-frame mismatch matters for
downstream 3D reconstruction. Pass
--rectify-output --rig-intrinsics-path data/rig_intrinsics.parquet to
audit_failures.py / eval_checkpoint.py to emit predictions in rectified
space instead. See the training repo for details.
Task: single-point keypoint detection on 4K underwater dive frames.
Training data: ~24,000 positive labels + hard-mined negatives across ~600 dives, dive-level splits (never frame-level). Two wavelengths (red laser and green laser). Multiple rigs (v1 is production; later rigs split off into their own models).
Input: 6 channels — chromaticity-normalized RGB (3) + wavelength (1) + Bayer-excess G / R (2). Native 4K resolution processed as overlapping 1024×1024 tiles with 256 px overlap.
Output: (x, y) pixel coordinate + presence sigmoid. See
predictions_with_meta.parquet schema in the training repo's HOW_TO_USE.md.
Architecture: segmentation_models_pytorch.Unet with ResNet-34 encoder
(production) or HRNet-w18 encoder (alternative), nearest-neighbor decoder
upsampling, 1-channel sigmoid heatmap output + per-tile presence head from
the encoder bottleneck.
Training: 50-epoch budget with patience-10 early stop; ~24h on 4× RTX 4500 Ada via DDP. Optimizer: AdamW, lr 3e-4, weight decay 1e-4, 1-epoch warmup. Loss: pixel-BCE-with-pos-weight (1000) for the heatmap + BCE for the presence head + optional line-consistency auxiliary loss.
Limitations:
rig_intrinsics.parquet you populate separately from the FishSense API.frames.parquet; those frames are
silently dropped during training and eval.License: BSD 3-Clause (same as the training code).
If you use this model, please cite the FishSense project:
@misc{ucsd_e4e_fishsense_laser_detector,
title = {FishSense Laser Detector},
author = {UCSD Engineers for Exploration},
year = {2026},
url = {https://huggingface.co/ucsde4e/fishsense-laser-detector}
}