Back to articles
AI Agents

DeskForge Builds Dense Supervision for Computer-Use Agents from Controllable Desktops

3 min read

Introduction

For a computer-use agent, understanding an instruction is only part of the problem. The agent must also identify the correct target on a crowded screen, where several applications, overlapping windows, and visually similar controls may appear at once. DeskForge presents a data-generation approach designed for this setting. Rather than relying only on static screenshots, it provides a controllable desktop environment that can vary application states, execute actions, and record what happened afterward.

What DeskForge provides

  • Controlled variation over real applications. The environment composes and explores real desktop applications while changing application state, content, window layout, visual appearance, and resolution. This preserves the complexity of practical software interfaces while making difficult conditions easier to generate systematically.
  • Dense multimodal annotations. DeskForge fuses screenshots with accessibility trees and window geometry. The result is a detailed description of interface elements and their locations, instead of supervision limited to a single final click point.
  • Action outcomes as supervision. Each executed action is recorded together with its result. This connects what the model sees with whether an attempted interaction had the intended effect.
  • A large corpus. DeskForge-1M contains 1.2 million annotated desktop observations and 159.7 million element instances. The authors use 200,000 grounding examples from this corpus to fine-tune four vision-language models.
  • Improvements beyond the training environment. According to the paper, all four models improve on held-out desktop conditions and on five external GUI grounding benchmarks. Qwen3.5-4B gains 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G.

Why it matters

The main contribution is not simply a larger dataset. It is an attempt to make the difficult parts of desktop interaction controllable. Manually collected demonstrations and isolated screenshots are unlikely to cover all combinations of occlusion, layout changes, resolution differences, and competing controls. By manipulating these factors within an environment that also exposes structural information, DeskForge can produce training examples targeted at the failure modes of GUI agents.

The reported effects also reach beyond local target selection. With a fixed planner, models fine-tuned on DeskForge data completed more tasks in the long-horizon WebArena-Infinity and OpenApps evaluations. For Qwen3.5-4B, the number of solved tasks rose from 31 to 50 out of 119 on WebArena-Infinity, and from 3 to 15 out of 100 on OpenApps. These results suggest that better grounding at individual steps can accumulate into more reliable end-to-end execution.

The findings should still be read as evidence for the proposed data and training approach, not as a claim that desktop agents are fully reliable. Changes in application versions, unexpected interface states, and cross-application workflows can create distribution shifts that a generated environment may not capture. The release of the framework, dataset, and fine-tuned model nevertheless gives future work a basis for reproducing the results and testing alternative methods for training computer-use agents.

arXiv

Comments

Checking sign-in status...

Loading comments...

Related articles