Downloads · 30 days
112
100% of all-time downloads
cklxx/laya-browser
laya-browser is a feature extraction model from cklxx. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
112
100% of all-time downloads
All-time downloads
112
Public
Parameters
322M
4.7 GB on disk
Likes
25
Trending 14
Click a slice to open those files.
.safetensors644 MB · 94%
From the Hugging Face model README

laya (convaiinnovations/laya) is a non-autoregressive "System 1" decision model:
one bidirectional encoder pass answers several typed questions (choice / score / noul) with calibrated probabilities, no text
generation. This repo fine-tunes it into the decision head of browser-use/jev-ultrafast,
whose /v1/systemone request is exactly laya's predict(state, questions): every step, one forward pass (~22 ms at full GPU clock) picks the operation
(CLICK / TYPE_TEXT / SELECT / PRESS_ENTER / SCROLL / DONE …) and its target element. Trained and evaluated locally on one RTX 4070
Ti SUPER (16 GB), no paid API; a local Qwen3-8B-AWQ only writes the text that TYPE_TEXT types and picks dropdown values.
One model, at the repo root (v19s). Earlier checkpoints are in the commit history. This repo is updated only when a new version is clearly better.
All numbers below were measured with the harness in code/jev-ultrafast.patch; the v17s column was measured with the harness it
shipped with, so part of the gain is the harness (see "What changed").
| evaluation | v17s | v19s |
|---|---|---|
| Suite C: 27 multi-step tasks on held-out real sites (search + filters + sort + open), ×2 | 2/54 | 11–14/54 (two ×2 runs; single ×1 runs: 5–7/27) |
| Suite B: 18 tasks on 18 held-out sites, ×3 | 54/54 | 54/54 |
| Suite A: 16 original real-site tasks, ×3 | 41/48 | 39/48 |
| webgym held-out synthetic forms, 7 kinds × 10 (flight, hotel, shop, car, restaurant, signup, filter) | 15/70 | 30/70 |
| held-out decisions on sites unseen in training (WebChain): click operation / target top-1 | — | 0.94 / 0.48 |
books-open-book (0/3), gained nothing else; suite B unchanged.Latency. One decision is 22–27 ms on an RTX 4070 Ti SUPER while the GPU is at full clock (requests back to back).
An agent waits seconds between decisions for pages to load, and the GPU drops to idle clocks (P5/P8, 210–850 MHz) within a
second or two: the next decision then takes 75–270 ms (the demo above shows these live numbers). Locking the minimum SM
clock removes that — sudo nvidia-smi -lgc 2100,3135 (undo: sudo nvidia-smi -rgc; costs some idle power). The first request
of each input-length bucket also compiles a kernel once (~100–500 ms); warm up with a few requests after starting the server.
transformers (the model ships its own code; only torch + transformers are needed):
from transformers import AutoModel
m = AutoModel.from_pretrained("cklxx/laya-browser", trust_remote_code=True).to("cuda") # CPU works too
page = {"url": ..., "title": ..., "text": ..., # the page: interactive elements + visible text
"actions": [{"id": "e1", "kind": "fill", "node": 1, "label": "Search packages", "role": "searchbox"},
{"id": "e2", "kind": "click", "node": 2, "label": "argo", "role": "link"}, ...]}
d = m.decide(page, goal="Search packages for 'json' and open the package 'argo'.", history=[])
d["operation"], d["action"], d["confidence"] # "CLICK", {"id": "e2", ...}, 0.88
decide builds the request exactly in the format the model was trained on (instructions, compact option strings, form-field
summary, 1200 chars of page text, split of choices wider than 60) and runs one forward pass. m.systemone(body) answers a
jev-ultrafast /v1/systemone request body. The same checkpoint also loads with the upstream laya package
(laya.load("cklxx/laya-browser")); laya_browser.py in this repo wraps it the same way and adds a server:
As a TypeSafe replacement for browser-use/jev-ultrafast:
python laya_browser.py serve --port 8791 # --model <local dir> to use a downloaded copy
TYPESAFE_BASE_URL=http://127.0.0.1:8791 TYPESAFE_API_KEY=local <run jev-ultrafast as usual>
It answers jev's /v1/systemone requests identically to the evaluation server (checked on 21 real steps: same operation and
target on all 21). The harness improvements behind the suite numbers are in code/jev-ultrafast.patch (apply to
jev-ultrafast 1231850); the full evaluation / training setup is in code/ (uv sync --extra fast, code/verify.py).
Harness (code/jev-ultrafast.patch, all measured case by case on failures):
MM/DD/YYYY fields) and picks the value of a dropdown
the policy chose (ages, times, countries that differ by a digit);Data (v19s = mmBERT-base v17s continued, full fine-tune):
code/finetune/convert_webchain.py);results/v19s/). A 322M single-pass policy does not reliably read these goal distinctions.| version | date | change |
|---|---|---|
| v10s | 2026-09-21 | first usable model: format v3, 17–23 ms per step |
| v14s | 2026-09-26 | harness fixes, Enter / scroll / select data, NNetNav; suite B 100 % |
| v17s | 2026-09-27 | dropdown option names kept, real category links; suite A 85 %, webgym 47 % (3 kinds) |
| v19s | 2026-09-29 | format v5, WebChain real-site trajectories, webgym ×7 + DAgger, harness fixes; suite C 4 % → ~20–26 %, webgym 21 % → 43 % (7 kinds) |
model.safetensors, encoder/, tokenizer/, rl_agent_config.json the model (v19s; a laya checkpoint dir)
modeling_laya_browser.py, config.json transformers remote code (AutoModel + trust_remote_code)
laya_browser.py the same on top of the laya package, plus a TypeSafe-compatible server
code/ server, suites A/B/C, Online-Mind2Web runner + judge, finetune pipeline (webgym, WebChain / Go-Browse converters),
TileLang kernels, jev-ultrafast patch, verify.py
results/ suite JSONs, traces and logs behind the numbers above (results/v19s/, older versions in their folders)
assets/ demo video
Apache-2.0, same as laya. Mind2Web, NNetNav and WebChain are used under their own licenses for training only.