Downloads · 30 days
11
73% of all-time downloads
kageroumado/rocuronium-ui-detector
rocuronium-ui-detector is a object detection model from kageroumado. Use it when you need objects located in an image. It is set up for coreml. The card lists the license as apache-2.0.
A tiny, fast UI-element box detector for desktop screenshots, exported to Core ML for Apple silicon. It proposes bounding boxes for on-screen controls so a GUI agent can aim at an app whose accessibility tree is empty…
Downloads · 30 days
11
73% of all-time downloads
All-time downloads
15
Public
Repo size
5.4 MB
Likes
1
Public
Click a slice to open those files.
.bin5.2 MB · 97%
From the Hugging Face model README
A tiny, fast UI-element box detector for desktop screenshots, exported to Core ML for Apple silicon. It proposes bounding boxes for on-screen controls so a GUI agent can aim at an app whose accessibility tree is empty or lying — icon toolbars included.
Trained for Rocuronium, a macOS UI-automation tool that drives apps without taking the cursor. It is the detector tier of Rocuronium's vision-grounding cascade (accessibility → detector + OCR → VLM).
confidence [N, 1] + coordinates [N, 4] — non-max-suppressed boxes in
the standard Core ML / Vision VNRecognizedObjectObservation format.UIElement. The model answers "a control is here", not "this is a
button vs a checkbox". Role and label come from OCR text inside the box (and, in
Rocuronium, a later VLM caption pass). This is deliberate — see Taxonomy below.GroundCUA's category field is 8
coarse buckets with ~33% of boxes empty, and it cross-cuts visual type — there is no clean
signal to learn Toggle / Slider / Checkbox / Link / Icon from. A box-proposal detector paired
with OCR for the label is both what the data supports and what the downstream consumer
actually uses, so the model commits to proposing boxes well rather than guessing a role
badly. A 3-class variant (Text / Interactable / Icon) is a documented follow-up.
Trained on a stratified ~13K-image subset of GroundCUA spanning all 87 apps, yolo11n, 640px, 100 epochs on Apple MPS.
| Metric | Value |
|---|---|
| mAP@50 | 0.881 |
| mAP@50-95 | 0.479 |
| precision | 0.864 |
| recall | 0.851 |
Core ML package: 5.4 MB. Latency: ~8 ms/image mean on .cpuAndNeuralEngine
(Apple silicon), 5.7 ms min.
import Vision
import CoreML
let model = try VNCoreMLModel(for: MLModel(contentsOf: compiledModelURL))
let request = VNCoreMLRequest(model: model)
request.imageCropAndScaleOption = .scaleFill
try VNImageRequestHandler(cgImage: screenshot).perform([request])
let boxes = (request.results as? [VNRecognizedObjectObservation] ?? [])
.filter { $0.confidence >= 0.25 } // normalized boundingBox, bottom-left origin
Trained on ServiceNow/GroundCUA (Apache-2.0). The weights are released under Apache-2.0 — unlike existing GroundCUA-trained detectors, which inherit Ultralytics' AGPL-3.0 and cannot ship in a closed app. Ultralytics tooling was used only to train; the exported weights carry no Ultralytics code.