Skip to content

xwm

SciWorld-MPO

xwm/SciWorld-MPO

SciWorld-MPO is a reinforcement learning model from xwm. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.

This model is a fine-tuned version of Llama-3.1-8B-Instruct on the sciworld-metaplan-preference-pairs dataset. It achieves the following results on the evaluation set: - Loss: 1.5017 - Rewards/chosen: -3.8774 - Reward…

Downloads · 30 days

19

6% of all-time downloads

All-time downloads

319

Public

Parameters

8B

16.1 GB on disk

Likes

2

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors16.1 GB · 100%

At a glance

Task
Reinforcement Learning
Library
transformers
License
apache-2.0
Model type
llama
Access
Public
Created
Feb 17, 2025
Updated
Mar 9, 2025
SHA
09cb62fa

Base models

Task
Reinforcement Learning
Library
transformers
Type
llama
License
apache-2.0
Languages
en
Created
Feb 17, 2025
Updated
Mar 9, 2025