Downloads · 30 days
19
6% of all-time downloads
xwm/SciWorld-MPO
SciWorld-MPO is a reinforcement learning model from xwm. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This model is a fine-tuned version of Llama-3.1-8B-Instruct on the sciworld-metaplan-preference-pairs dataset. It achieves the following results on the evaluation set: - Loss: 1.5017 - Rewards/chosen: -3.8774 - Reward…
Downloads · 30 days
19
6% of all-time downloads
All-time downloads
319
Public
Parameters
8B
16.1 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of Llama-3.1-8B-Instruct on the sciworld-metaplan-preference-pairs dataset. It achieves the following results on the evaluation set:
See the original paper for more details: MPO: Boosting LLM Agents with Meta Plan Optimization.
Code: https://github.com/WeiminXiong/MPO
This model uses Meta Plan Optimization (MPO) to improve the planning capabilities of LLM agents. It leverages high-level general guidance through meta plans and enables continuous optimization based on feedback from the agent's task execution. It achieves state-of-the-art performance on ALFWorld and SciWorld, with an average accuracy of 83.1.
More information needed
The model was trained on the sciworld-metaplan-preference-pairs dataset, part of the Meta_Plan_Optimization dataset.
The following hyperparameters were used during training: