OSWorld-Science Tests Whether AI Agents Can Actually Use Scientific Software
Using a computer in a research setting involves much more than recognizing a screen and clicking the right control. An agent must understand specialized interfaces, manipulate scientific objects, configure analyses or simulations, and produce an output that can be checked independently. OSWorld-Science is designed to evaluate this broader capability.
A benchmark built around research workflows
The benchmark contains 146 tasks and covers 12 vision-language models across multiple scientific domains and software configurations. Its examples include molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation. These tasks move beyond generic desktop interaction by asking whether an agent can complete a meaningful chain of scientific actions.
Task proposals come from experts and are refined through an iterative human–AI co-design process. They are then selected according to scientific value and difficulty. This approach is intended to reduce the gap between a task that merely resembles research software usage and one that actually captures a research objective.
Artifact-based evaluation
A central feature of OSWorld-Science is its use of execution-based, task-specific evaluators. Rather than relying only on screenshots or interaction traces, the evaluators inspect application states and the artifacts produced by the agent. Depending on the task, these artifacts may include:
- molecular structures;
- pathology segmentation masks;
- statistical plots;
- numerical results from calculations or simulations.
The benchmark can also award partial credit when an agent completes only part of a workflow. This makes it possible to distinguish between failures in understanding the instruction, operating the software, and producing the final scientific result. It also avoids reducing every imperfect trajectory to a simple pass-or-fail label.
The harness is part of the experiment
The project provides an agent harness that combines model adapters, interaction-loop control, and trajectory logging. This allows researchers to compare not only models, but also the systems surrounding those models. Planning policy, context handling, stopping decisions, and the organization of visual and command-line actions can all influence the final outcome.
The reported results suggest that state-of-the-art VLMs still face substantial difficulties on important scientific tasks, even when paired with a strong harness. The study also examines factors such as multilingual inputs, reasoning effort, and context length, offering directions for future work. The available material does not provide detailed per-model scores or task-level failure analyses, so it cannot support claims about which model is best in a particular discipline.
Why it matters
OSWorld-Science connects expert-defined scientific goals with software execution and verifiable outputs. That connection is important because a research agent should not be evaluated solely by whether it can navigate an interface. It should also be judged by whether the resulting structure, mask, plot, or number satisfies the scientific objective.
The benchmark may therefore become useful for studying both agent capabilities and harness design. Its longer-term value will depend on whether the community adds more realistic tasks and maintains reliable artifact evaluators. For now, it offers a clearer framework for measuring the gap between general computer use and dependable scientific assistance.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...