Downloads · 30 days
39
41% of all-time downloads
OS-Copilot/OS-Shepherd-35B-A3B
OS-Shepherd-35B-A3B is a image-text-to-text model from OS-Copilot. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
OS-Shepherd-35B-A3B is an open multimodal reward model for judging computer-use agent trajectories. Given a task instruction, screenshots, and the agent's reasoning and actions, it determines whether the task was comp…
Downloads · 30 days
39
41% of all-time downloads
All-time downloads
94
Public
Parameters
35.1B
70.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors70.2 GB · 100%
From the Hugging Face model README
OS-Shepherd-35B-A3B is an open multimodal reward model for judging computer-use agent trajectories. Given a task instruction, screenshots, and the agent's reasoning and actions, it determines whether the task was completed and returns a reasoned SUCCESS or FAIL verdict.
It is fine-tuned from Qwen3.5-35B-A3B on OS-Shepherd-100K using the same SFT and GRPO recipe as OS-Shepherd-9B. The RL stage focuses on reducing false-success judgments.
| Benchmark | Accuracy | Fail recall |
|---|---|---|
| OSReward | 85.6 | 86.2 |
| OSReward-Hard | 62.7 | 60.1 |
Results use the fixed judging protocol described in the OSReward paper.
Use the canonical prompt and trajectory format from the OSReward repository. A recent Transformers, vLLM, or SGLang version with Qwen3.5 multimodal support is required.
This model is intended for trajectory evaluation, data filtering, and reward-model research. It is not a computer-control policy and may still miss fine-grained visual failures, especially on hard cases.
Apache License 2.0. See LICENSE.
@article{sun2026osreward,
title={OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models},
author={Sun, Qiushi and others},
journal={arXiv preprint arXiv:2607.28609},
year={2026}
}