Back to articles
World Models

AutoGUIWorld Uses Image Generators to Simulate GUI Environments

3 min read

Introduction

A capable computer-use agent must do more than locate buttons. It needs to understand how an action changes the interface, preserve state across multiple steps and recover from the evolving context of a task. Building such competence traditionally requires running real applications and recording interaction traces. That approach is costly: applications must be installed and configured, while specialized tools and different operating systems add further engineering overhead.

AutoGUIWorld proposes to move part of this process into a synthetic visual environment. Its central idea is to use an image generator as a visual world model, producing the next screen after a planned GUI action without actually executing the corresponding software.

How the framework works

  • Structured scene initialization. The system samples initial GUI scenes from specifications covering operating-system context, visual appearance and interface state. This gives the data pipeline control over the conditions under which a task begins.
  • Planner-driven task construction. A planner proposes tasks and breaks them into atomic actions. It also describes the intended visual consequence of each action, linking an operation to the state change that should follow.
  • Iterative screenshot editing. Instead of launching an application after every action, an image generator edits the current screenshot to produce the next observation. Repeating this process yields multi-step trajectories with changing visual states.
  • Grounding and filtering. Action grounding and transition-level quality checks are used to remove weak examples and retain spatially annotated step-level samples. The reported dataset contains 79,266 samples spanning Ubuntu, Windows, macOS and Chrome.

Results and implications

The authors fine-tuned Qwen3.5-35B-A3B on AutoGUIWorld trajectories. On OSWorld, the mean task score increased from 33.0% to 40.8%. On ScienceBoard, the task success rate rose from 14.0% to 32.2%. These results suggest that synthetic visual transitions can provide useful supervision for the relationship between an action, the resulting screen and the broader task state.

The broader contribution is a change in the economics of GUI data generation. Instead of treating every application as a runtime dependency, a data creator can specify a scene and synthesize a plausible visual transition. This may make it easier to expand coverage across operating systems, interface layouts and workflows that are difficult to deploy at scale.

What remains uncertain

A visually plausible screenshot is not the same as a faithful software environment. Image generation may reproduce the appearance of a menu opening while missing hidden application logic, permission errors, loading behavior or other irregular feedback. The usefulness of synthetic trajectories therefore depends on how well they transfer to executable environments. AutoGUIWorld should be viewed as a complement to real interaction data and real-world evaluation, rather than a complete replacement. Calibrating generated transitions against actual software behavior will be important for future progress.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles