Downloads · 30 days
154
47% of all-time downloads
lumimate/PhoneUIAnchor-829M
PhoneUIAnchor-829M is a image-text-to-text model from lumimate. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
PhoneUIAnchor-829M is a vision-language model for locating elements in mobile and web interfaces. Given a screenshot and a natural-language description, it returns the normalized center point of the target element. It…
Downloads · 30 days
154
47% of all-time downloads
All-time downloads
325
Public
Parameters
829M
3.3 GB on disk
Likes
14
Public
Click a slice to open those files.
.safetensors3.3 GB · 100%
From the Hugging Face model README
PhoneUIAnchor-829M is a vision-language model for locating elements in mobile and web interfaces. Given a screenshot and a natural-language description, it returns the normalized center point of the target element. It can be used with GUI agents, test automation, accessibility tools, and similar visual workflows.
This repository contains a complete, standalone checkpoint that can be loaded directly with Transformers without additional model files.
| Property | Value |
|---|---|
| Architecture | Florence-2-large |
| Task | GUI element grounding |
| Parameters | 829M |
| Output contract | <loc_x>,<loc_y> |
| Weight format | Safetensors |
| Recommended precision | BF16 |
| Tested stack | Python 3.10, PyTorch 2.4.1, CUDA 12.1 |
pip install -r requirements.txt
The packaged Florence-2 implementation uses custom model code. Review the
included Python files and load with trust_remote_code=True.
import re
import torch
from PIL import Image
from transformers import AutoModelForCausalLM, AutoProcessor
model_id = "lumimate/PhoneUIAnchor-829M"
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
attn_implementation="sdpa",
).cuda().eval()
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
image = Image.open("ui_screenshot.png").convert("RGB")
prompt = (
'Where is the "Settings" element? '
'(Output the center coordinates of the target)'
)
inputs = processor(images=image, text=prompt, return_tensors="pt").to(
"cuda", dtype=torch.bfloat16
)
with torch.inference_mode():
output_ids = model.generate(
**inputs,
do_sample=False,
max_new_tokens=16,
)
text = processor.tokenizer.batch_decode(
output_ids, skip_special_tokens=False
)[0]
match = re.search(r"<loc_(\d+)>,<loc_(\d+)>", text)
point = tuple(map(int, match.groups())) if match else None
x_px = point[0] / 999 * image.width
y_px = point[1] / 999 * image.height
print({"normalized": point, "pixels": (x_px, y_px)})
A command-line implementation is included in inference_example.py.
PhoneUIAnchor-829M is intended for research and product prototyping in GUI perception, agentic interaction, automated testing, and assistive interfaces. Predictions should be validated before high-impact or irreversible automation.
PhoneUIAnchor-829M uses the Florence-2 architecture and is distributed under the
terms provided in LICENSE.