Downloads · 30 days
331
100% of all-time downloads
RUC-AIBOX/AEWM
AEWM is a text generation model from RUC-AIBOX. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
AEWM is a world model that judges and edits an agent's decisions, rather than predicting tool responses. Built on Qwen3.5-35B-A3B, it combines Action Judge (AJ) and State Revision (SR) to improve long-horizon reasonin…
Downloads · 30 days
331
100% of all-time downloads
All-time downloads
331
Public
Parameters
35.1B
70.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors70.2 GB · 100%
From the Hugging Face model README
AEWM is a world model that judges and edits an agent's decisions, rather than predicting tool responses. Built on Qwen3.5-35B-A3B, it combines Action Judge (AJ) and State Revision (SR) to improve long-horizon reasoning and tool use across Search, Terminal, and Software Engineering (SWE). Its inference framework, EditAct, integrates these capabilities into an agent's interaction with the real environment.
<!-- The arXiv link is intentionally empty until the paper is available. --> <p align="center"> <a href=""><img src="https://img.shields.io/badge/Paper-arXiv-B31B1B?logo=arxiv&logoColor=white" alt="Paper (arXiv link forthcoming)"></a> <a href="https://huggingface.co/RUC-AIBOX/AEWM"><img src="https://img.shields.io/badge/🤗_Hugging_Face-AEWM-FFD21E" alt="AEWM model on Hugging Face"></a> <a href="https://huggingface.co/datasets/RUC-AIBOX/AEWM-Action-Judge-Bench"><img src="https://img.shields.io/badge/🤗_Hugging_Face-Benchmark-FFD21E" alt="Action Judge Benchmark on Hugging Face"></a> <a href="https://github.com/RUCAIBox/Agent-Editing-World-Model"><img src="https://img.shields.io/badge/Code-GitHub-181717?logo=github&logoColor=white" alt="Code on GitHub"></a> </p>| Property | Description |
|---|---|
| Base model | Qwen3.5-35B-A3B |
| Capabilities | Action Judge and State Revision in one checkpoint |
| Domains | Search, Terminal, and Software Engineering |
| Training | Cross-domain mid-training followed by supervised fine-tuning |
| Release format | Full model weights in Hugging Face Safetensors format |
| Inference framework | EditAct |
This checkpoint serves as the world model, alongside a separate agent that proposes reasoning and actions.
Long-horizon agents can suffer from task-state contamination: unsupported assumptions become accepted facts, outdated plans persist, and partial progress is mistaken for completion. Subsequent actions may appear locally reasonable while reinforcing an incorrect understanding of the task.
AEWM addresses this problem by modeling how an agent's reasoning and actions shape task progress. Given the observed history and a candidate decision, it judges the decision's contribution and, when necessary, supplies a revised reasoning-action continuation.
AJ classifies a proposed decision into one of three action types:
This judgment determines whether to retain the proposal or intervene. At inference time, AJ uses only information available before the candidate action is executed.
For a noisy decision, SR produces revised reasoning and action from the same observed history. It directly edits the continuation that enters the agent's execution loop, rather than merely adding a critique or asking the agent to try again.
EditAct runs four steps:

Method overview from the paper: learning Action Judge and State Revision, integrating them into EditAct, and transferring guided behavior back into an agent through AEWM-RFT.
AEWM is trained in two stages. Mid-training uses approximately 52B tokens of cross-domain agent trajectories, AJ supervision, and SR supervision. Supervised fine-tuning uses 120K curated examples: 60K AJ and 60K SR examples, with 40K examples per domain.
AJ supervision labels decisions using trajectory evidence and consistency checks. SR supervision pairs agent proposals with corrected reasoning-action continuations from the same history. Quality filtering checks annotation consistency and revision quality. The resulting checkpoint supports both judgment and revision across different tool interfaces.
AEWM-RFT is a separate use of the method: an agent is fine-tuned on verified EditAct trajectories to internalize useful editing behavior. The model in this repository is AEWM itself, used for online judgment and revision.
The AEWM Action Judge Benchmark contains 3,000 annotated decisions, with 1,000 each from Search, Terminal, and SWE. It evaluates decision classification before execution, reporting accuracy and macro-F1.

AEWM achieves 70.5% overall macro-F1, exceeding the strongest compared baseline, DeepSeek-V4-Pro (59.9%), by 10.6 percentage points. Its macro-F1 scores are 60.9% on Search, 72.1% on Terminal, and 77.8% on SWE. Overall metrics pool all 3,000 decisions.
The paper evaluates EditAct with Qwen3.5-4B, Qwen3.5-9B, and Qwen3.5-35B-A3B on BrowseComp, DeepSearchQA, Terminal-Bench 2.0, SWE-bench Pro, Doc2Repo, and NL2Repo.
| backbone | ReAct | Strongest baseline average | EditAct | Gain over strongest baseline |
|---|---|---|---|---|
| Qwen3.5-4B | 28.5 | 35.1 | 41.8 | +6.7 |
| Qwen3.5-9B | 34.5 | 38.9 | 44.1 | +5.2 |
| Qwen3.5-35B-A3B | 42.2 | 45.6 | 48.8 | +3.2 |
Paper results: mean score (%) across the six benchmarks. The strongest baseline is selected by its six-benchmark average among ReAct, step-level Best@3, and trajectory-level Best@3. Gains are absolute percentage points.
Serve this checkpoint with a backend that supports Qwen3.5 reasoning and native tool calls, then connect it to the EditAct framework. The agent and AEWM use separate OpenAI-compatible endpoints. Set the world-model endpoint as follows, matching WM_MODEL to the name configured by your serving backend:
export WM_MODEL="RUC-AIBOX/AEWM"
export WM_BASE_URL="http://localhost:30001/v1"
export WM_API_KEY="EMPTY"
EMPTY is only for an endpoint without authentication. Configure the agent and tools separately, following the repository's quick start. Use the provided domain-specific AJ/SR prompts and preserve reasoning and tool-call fields; a generic chat prompt is not equivalent to the evaluated workflow.
The framework provides editact_for_search, editact_terminal, and editact_for_swe scaffolds. Search requires search and webpage-reading services; Terminal and SWE require an execution sandbox. The standalone Action Judge evaluator instead classifies recorded decisions without executing tools.