Downloads · 30 days
23
13% of all-time downloads
osunlp/Dreamer-7B-Classifieds
Dreamer-7B-Classifieds is a image-text-to-text model from osunlp. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
WebDreamer is a planning framework that enables efficient and effective planning for real-world web agent tasks. Check our paper for more details. This work is a collaboration between OSUNLP and Orby AI.
Downloads · 30 days
23
13% of all-time downloads
All-time downloads
172
Public
Parameters
8.3B
16.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
WebDreamer is a planning framework that enables efficient and effective planning for real-world web agent tasks. Check our paper for more details. This work is a collaboration between OSUNLP and Orby AI.
root
|-- prompt: string
|-- image: binary
|-- response: string
|-- action: string
| Benchmark | Method | Success Rate |
|---|---|---|
| VisualWebArena | GPT-4o + Reactive | 17.6% |
| GPT-4o + Tree Search | 26.2% | |
| GPT-4o + WebDreamer | 23.6% (↑34.1%) | |
| Online-Mind2Web | GPT-4o + Reactive | 26.0% |
| GPT-4o + WebDreamer | 37.0% (↑42.3%) | |
| Mind2Web-live | GPT-4o + Reactive | 20.2% |
| GPT-4o + WebDreamer | 25.0% (↑23.8%) |
Compared to the reactive baselines, WebDreamer significantly improves performance by 34.1%, 42.3%, and 23.8% on VisualWebArena, Online-Mind2Web, and Mind2Web-live, respectively.
WebDreamer effectively explores the search space through simulations, which largely reduces the reliance on real-world interactions while maintaining robust performance.
vllm serve osunlp/Dreamer-7B --api-key token-abc123 --dtype float16
or
python -m vllm.entrypoints.openai.api_server --served-model-name osunlp/Dreamer-7B --model osunlp/Dreamer-7B --dtype float16
You can find more instruction about training and inference in Qwen2-VL's Official Repo.
Actually our model is quite robust to textual prompt so feel free to try various prompts which we didn't heavily explore.
def format_openai_template(description: str, base64_image):
return [
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {"url": f"data:image/jpeg;base64,{base64_image}"},
},
{
"type": "text",
"text": f"""
Below is current screenshot. Please describe what you would see after a {action_description}"""
},
],
},
]
messages = format_openai_template(description, base64_image)
completion = await client.chat.completions.create(
model=args.model_path,
messages=messages,
temperature=1.0
)
If you find this work useful, please consider citing our papers:
@article{Gu2024WebDreamer,
author = {Yu Gu and Kai Zhang and Yuting Ning and Boyuan Zheng and Boyu Gou and Tianci Xue and Cheng Chang and Sanjari Srivastava and Yanan Xie and Peng Qi and Huan Sun and Yu Su},
title = {Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents},
journal = {CoRR},
volume = {abs/2411.06559},
year = {2024},
url = {https://arxiv.org/abs/2411.06559},
eprinttype= {arXiv},
eprint = {2411.06559},
}