HybridCUA Teaches Computer-Use Agents to Orchestrate GUI and CLI
Introduction
Computer-use agents are increasingly expected to complete tasks inside real software environments rather than merely answer questions. Most current systems interact through a graphical user interface (GUI), clicking, typing, dragging, and reading the screen. This approach is broadly applicable, but it can be slow and vulnerable to layout changes or localization errors. A different strategy is to provide application-specific APIs or tools, yet building and maintaining those integrations requires considerable engineering effort and makes scaling across applications difficult.
HybridCUA proposes a middle path: an agent should be able to use both the GUI and the command-line interface (CLI), and should learn how to coordinate them during a task.
The central challenge is tool orchestration
The contribution is not simply to give a model access to a terminal. The harder problem is deciding when the terminal is useful, what command should be executed, and how its result should be connected to subsequent GUI actions. GUI interaction offers generality across applications and is useful when the agent must inspect visual state or manipulate elements directly. CLI operations can be much more efficient for tasks such as file handling, batch processing, or other operations naturally expressed as shell commands.
To teach this behavior, the authors construct three types of trajectories: GUI-only, CLI-only, and interleaved GUI-and-CLI trajectories. The resulting HybridCUA-8K dataset contains 5K hybrid trajectories and 3K verified RLVR tasks. This design gives the model more than a collection of tool calls. It exposes alternative execution patterns and provides examples of how an agent can move between visual interaction and command execution.
Training has two stages. Supervised fine-tuning first teaches the action patterns represented in the constructed trajectories. The second stage uses online reinforcement learning with CLI-aware rewards. These rewards are intended to encourage selective and reliable command-line use, rather than maximizing the number of terminal actions. In principle, that distinction matters: an agent that invokes the shell indiscriminately may introduce new errors instead of reducing interaction cost.
Results and implications
HybridCUA-9B achieves 53.6% accuracy on OSWorld, an improvement of 14.8 percentage points over the base model. It also improves performance on WindowsAgentArena by 4.0 percentage points. Based on the reported results, the experiments support the idea that coordinated GUI and CLI behavior can improve computer-use performance and transfer across platforms.
The broader lesson is that tool selection should be treated as a core capability of an agent, not as an afterthought. A general-purpose computer agent may need visual interaction for some steps and concise shell operations for others. Hybrid trajectories and tool-aware rewards offer one way to train this decision-making behavior.
The available material does not provide detailed task breakdowns, failure analyses, or ablations for individual CLI strategies. Therefore, the reported gains should be read as evidence for the overall approach, while its robustness across more applications and operating environments remains an open question.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...