Downloads · 30 days
0
abedinia/laya-web-agent
laya-web-agent is a machine learning model from abedinia. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for laya. The card lists the license as apache-2.0.
A fast decision model for browser agents. Given a user's goal and the current page, it picks the next step in a single forward pass, without generating any text:
Downloads · 30 days
0
Access
Public
Updated Sep 28, 2026
Parameters
421M
6.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.7 GB · 100%
From the Hugging Face model README
A fast decision model for browser agents. Given a user's goal and the current page, it picks the next step in a single forward pass, without generating any text:
CLICK, TYPE_TEXT, SELECT, SCROLL_DOWN, WAIT, DONE or BLOCKEDThe input is the page state a browser agent sends: title, visible text, a numbered list of the controls on screen with their current values, and the recent actions.
| Architecture | Laya decision model on a ModernBERT-large encoder (28 layers, 1024 hidden) |
| Parameters | 421M |
| Context | 2048 tokens (up to 512 for the question and its options) |
| Output | typed answers with probabilities, one forward pass |
import laya
from huggingface_hub import snapshot_download
agent = laya.load(snapshot_download("abedinia/laya-web-agent"))
The input format is jev_ultrafast's: the state plus typed operation and target questions.
rl_agent_config.json holds the lengths the model was trained with (max_len 2048,
head_max_len 512). Keep them. With shorter lengths, field values and history get cut off.
For each step, the top 25 candidate elements come from the MindAct DeBERTa ranker (its scores ship with the dataset), listed in page order. A step counts as a miss if the ranker drops the correct element, or if the step couldn't be converted (about 1%).
| split | steps | element | operation | step success | step success, macro per task |
|---|---|---|---|---|---|
| cross-task (new tasks, known sites) | 2,094 | 35.0% | 85.8% | 26.1% | 29.1% |
| cross-website (new sites) | 1,373 | 27.1% | 82.1% | 18.6% | 21.4% |
| cross-domain (new domains) | 5,911 | 27.8% | 84.5% | 19.4% | 21.7% |
500 held-out episodes (2,381 decisions) of 17 page types, generated from seeds training never used and played in real Chrome:
| per step, both questions | 98.5% |
| operation | 97.8% |
| target | 99.4% |
| episodes with every step right | 87.7% (436 of 497) |
It's at 100% on checkboxes, wizards, contact messages, already-completed pages and impossible tasks (BLOCKED, 12/12). The weakest page types are carts (86.5% on operation) and recovering after a mistyped form value (90.2%). Most remaining mistakes are CLICK-or-SCROLL_DOWN calls when a control sits at the very bottom of the screen.
Data: 33,438 decisions.
Labels: CLICK, TYPE_TEXT, SELECT, SCROLL_DOWN (1,552), WAIT (421), DONE (5,474) and BLOCKED (143).
Setup: