GUI-HARVEST Helps GUI Agents Improve Through Execution Evidence
Introduction
A GUI agent is more than a vision-language model that can read a screen and emit clicks. Its surrounding execution harness determines how observations are assembled, how outputs become actions, and how the system verifies success, recovers from errors, or decides to stop. GUI-HARVEST focuses on improving this layer rather than updating the model itself. The backbone remains frozen while the harness evolves from evidence collected during execution.
How the approach works
GUI failures are difficult to diagnose because the model’s intent and the interface’s response can diverge. A model may select a plausible action, yet a coordinate error, delayed response, or unexpected application state can lead to a different screen. GUI-HARVEST grounds its diagnosis in observable transitions rather than relying only on a final success label.
Its workflow has three main elements:
- Align intent, actions, and visual effects. Model outputs and executed actions are matched with screenshots captured before and after the action. This ties a suspected problem to a concrete interface transition.
- Use repeated runs as joint evidence. The same task may produce different outcomes because of timing, state, or execution variability. Comparing multiple runs helps identify behavioral differences that actually correlate with success or failure.
- Convert recurring findings into bounded edits. Verified findings from different tasks are consolidated into recurring failure patterns. These patterns are mapped to constrained source-code changes, while the expected behavioral effect is recorded before evaluation. The system then checks both the predicted effect and task-level performance.
The result is a closed loop of observation, diagnosis, modification, and verification. This is more structured than asking a language model to invent a new prompt or freely rewrite an agent scaffold.
Results and implications
On OSWorld-Verified, the study covers six general-purpose, GUI-specialized, and proprietary backbones. GUI-HARVEST achieves the highest score in every model-and-budget comparison shown in the supplied material. The harness is optimized at a 15-step setting and then kept frozen for evaluation at 15, 50, and 100 steps. Qwen3-VL-32B-Instruct improves by 12.33 percentage points over its initial harness at 15 steps, while Gemini 3.1 Pro reaches 79.14% at 100 steps.
The broader message is that GUI-agent progress does not have to come solely from larger or better-trained models. Observation assembly, action execution, verification, recovery, and termination can each become a system bottleneck. Harness evolution offers a way to improve these components without changing backbone weights, while also making engineering changes easier to inspect and evaluate.
There are still important constraints. The approach depends on the quality of screenshots and traces, the reliability of failure attribution, and the ability of code edits to avoid regressions. Cross-task reuse may also introduce new risks when a fix works in one interface but is not appropriate elsewhere. Scaling evidence-driven harness evolution to more open-ended desktop environments will therefore require stronger safeguards and broader validation.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...