Results
Table with columns: Benchmark, Accuracy, Fail recall| Benchmark | Accuracy | Fail recall |
|---|
| OSReward | 85.6 | 86.2 |
| OSReward-Hard | 62.7 | 60.1 |
Results use the fixed judging protocol described in the OSReward paper.
Usage
Use the canonical prompt and trajectory format from the OSReward repository. A recent Transformers, vLLM, or SGLang version with Qwen3.5 multimodal support is required.
This model is intended for trajectory evaluation, data filtering, and reward-model research. It is not a computer-control policy and may still miss fine-grained visual failures, especially on hard cases.
License
Apache License 2.0. See LICENSE.
Citation
@article{sun2026osreward,
title={OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models},
author={Sun, Qiushi and others},
journal={arXiv preprint arXiv:2607.28609},
year={2026}
}