OSWorld-Pro Moves Computer-Use Evaluation Beyond the Final Result
Introduction
Computer-use agents (CUAs) operate keyboards, mice, and graphical interfaces to complete multi-step tasks. Existing evaluations often reduce the result to a final-state question: was the requested file created, was the application configured correctly, or did the page end in the expected state? This approach is convenient, but it hides the path taken by the agent and makes failures difficult to diagnose.
OSWorld-Pro, introduced by an NVIDIA research team, shifts attention from the final deliverable to the sequence of actions that produces it. Instead of assigning only a task-level success or failure label, the benchmark examines whether an agent completes a series of dependent subgoals and where its execution begins to diverge.
Key points
- Tasks are decomposed into procedural steps. OSWorld-Pro contains more than 300 tasks and over 2,800 subgoals. This structure makes it possible to track progress across an interaction rather than treating the entire trajectory as a black box.
- The benchmark is grounded in human annotations. Its evaluation is based on more than 67,000 human annotations. Human-aligned LLM judges are then used to assess whether individual subgoals have been fulfilled, preserving a finer-grained view than a single functional verifier can provide.
- Failure modes become more actionable. The study highlights behaviors such as actions unrelated to the current subgoal and mistakes in click-based interaction. A keyboard-entry failure calls for a different intervention from inaccurate visual targeting, yet both would look identical under a simple final-state metric.
- Process evaluation is harder. Claude Opus 5 reportedly reaches 75.7% on OSWorld-Pro, compared with 83.4% on OSWorld. The scores should not be treated as perfectly interchangeable, but the gap illustrates how much instability can remain hidden when only the endpoint is checked.
Why it matters
For developers, a procedural benchmark offers better debugging signals. Repeated keyboard mistakes may point to weaknesses in action modeling, state verification, or recovery. Misplaced clicks could require improvements in visual grounding, coordinate mapping, or post-click confirmation. Frequent irrelevant actions may instead indicate problems with planning, context tracking, or stopping behavior.
This distinction is important because an agent can sometimes reach the right outcome through unnecessary or fragile actions. Such behavior may look successful in a final-state test but fail when the interface changes, the task becomes longer, or an intermediate error requires recovery. Tracking subgoals exposes whether the agent is consistently competent or merely succeeds intermittently.
LLM-based judging still needs careful validation, since its usefulness depends on alignment with human assessments. Even so, OSWorld-Pro represents a meaningful change in evaluation philosophy. For computer-use agents, the central question is no longer only whether the final artifact is correct, but also whether the agent followed a reliable, efficient, and understandable path to produce it.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...