AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

arXiv · AI, language, vision and robotics · article · Sep 21, 2026 · UTC

Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agents fail in various tasks, obfuscating critical insight for subsequent improvement. For instance, agents that err during keyboard inputs would require a different mitigation strategy from those that fail to precisely provide click-based inputs on the graphical UI. We introduce OSWorld-Pro: a set of over 300 tasks contai

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T06:11:12.848Z. This is not the publication date.