Downloads · 30 days
14
100% of all-time downloads
benjbritton/chactun-multiclass-detector
chactun-multiclass-detector is a image segmentation model from benjbritton. Use it for the image segmentation task on the model card, and read the license before you ship it in a product. It is set up for detectron2. The card lists the license as cc-by-4.0.
Mask R-CNN R50-FPN that detects three classes of ancient Maya feature — building, platform, aguada — in airborne lidar relief visualizations of the central Yucatán Peninsula.
Downloads · 30 days
14
100% of all-time downloads
All-time downloads
14
Public
Repo size
351 MB
Likes
0
Public
Click a slice to open those files.
.pth351 MB · 100%
From the Hugging Face model README
Mask R-CNN R50-FPN that detects three classes of ancient Maya feature — building, platform, aguada — in airborne lidar relief visualizations of the central Yucatán Peninsula.
Trained on the Chactún ML-ready dataset (Kokalj et al. 2023), on the ECML PKDD 2021 discovery challenge's own training split. Full experimental record, including four claims that later measurement refuted, in benjbritton/geoai-detection.
Candidate generation, not a census. The catalogue this model produces is a probabilistic candidate layer for landscape survey. Individual detections should be treated accordingly — see Limitations.
The model was trained on a specific rendering, and it is only as portable as that rendering is reproducible. Three bands at 0.5 m:
| band | visualization | stretch |
|---|---|---|
| 1 | sky-view factor, 5 m radius, 16 directions | linear 0.7 – 1.0 |
| 2 | positive openness, same geometry | linear 68° – 93° |
| 3 | slope | inverted greyscale, linear 0° – 50° |
These are the general-terrain settings from Table 3 of Kokalj et al. 2023. Applying a different stretch to the same physical quantities produces a representation the model has not seen, and performance degrades in ways that are not obvious from the output. Matching band statistics is not the same as applying the same stretch function — that was measured, and it was worse than doing nothing.
Evaluated on the challenge's held-out split (tiles 1765–2093). Single seed.
| metric | value |
|---|---|
| segm AP | 44.63 |
| segm AP50 | 69.31 |
| segm AP75 | 49.17 |
| building / platform / aguada | 42.39 / 55.25 / 36.24 |
| semantic IoU (challenge convention) | 0.794 |
For context, the published leaderboard field on that split runs 0.811–0.834 semantic IoU. Those entries are five-fold ensembles with pseudo-labeling and test-time augmentation; this is a single model. A semantic-segmentation baseline trained at matched compute scores 0.809, which locates most of the difference in the instance-versus-semantic formulation rather than in training effort.
Operating point matters more than AP for survey use. Pooled over all 2,094 tiles, at score 0.05 the detector recalls 92% of annotated structures at 188 false positives per km²; at 0.50, 75% recall at 32 FP/km². Choose against the cost of a field visit, not against the metric.
Mask R-CNN R50-FPN, COCO-pretrained, detectron2. The one intervention that moved the problem was D4 augmentation — the eight symmetries of a square — worth +4.16 AP over the control across five-fold cross-validation, 95% CI [+2.70, +5.61], positive on every fold, and replicated at +4.17 on this split.
It is valid here because these three bands are computed isotropically, so a rotation is label-preserving. It would not be valid on hillshade, where a fixed illumination azimuth is baked into the pixels.
Four model-side interventions were tested and came in within noise: shifted anchors, a cascade head, 960 px input, and repeat-factor oversampling. Their confidence intervals all contain zero.
It was measured, on this model's own validation split, so the comparison is like for like: same 329 tiles, same annotations, bbox AP through the same COCOeval path.
| bbox AP | |
|---|---|
| this model (arm D) | 42.33 |
| Grounding DINO base, best of four prompt formulations | 3.08 |
| OWLv2 base, best of four | 2.64 |
Neither zero-shot result is noise — both sit well above a size-matched random-box null — but a factor of fourteen is not a gap that prompt engineering closes. On overhead optical imagery the same checkpoints do considerably better, reaching about a quarter of their supervised counterpart rather than a fourteenth, which suggests the shortfall is about the imagery rather than about these particular structures.
Two practical findings from that work transfer to anyone prompting these models on terrain visualizations. Describing shape beats naming the class, replicated across both model families. And confidence is anti-correlated with correctness off-distribution: Grounding DINO's highest-scoring boxes were whole-tile in every arm, so no score threshold recovers a usable detector, though an area prior removes the degenerate band trivially.
Method, figures and code: https://github.com/benjbritton/geoai-detection
Weights released CC BY 4.0, matching the training data.
Kokalj, Ž., Somrak, M., et al. (2023). Machine learning-ready remote sensing data for Maya archaeology. Scientific Data 10:558. https://doi.org/10.1038/s41597-023-02455-x
If you use this model, cite the dataset. The code is MIT; the weights inherit CC BY 4.0 from the data.
| file | |
|---|---|
model_final.pth | detectron2 checkpoint, 351 MB |
config.yaml | the training config, layered on the model-zoo base |
inference.py | minimal example: load, run on a 3-band tile, write GeoJSON |
The GeoJSON follows the coordinate space of the input and records which it
used, as coordinate_space. A georeferenced raster gives map coordinates
in its own CRS. The Chactun tiles are ungeoreferenced by design, so they
give pixel coordinates, and the script says so when it runs.