Back to articles
AI Agents

CompoWorld Trains Agents to Complete Cross-Service Workflows

3 min read

Introduction

Many real-world agent tasks cross application boundaries. An agent may need to retrieve information from one service, transform it, and then use the result in a calendar, project-management, or automation tool. Training environments that capture this kind of dependency are difficult to build at scale. CompoWorld addresses the problem with a compositional approach to environment generation.

How the system works

Rather than attempting to create an entirely new environment for every task, CompoWorld first builds a library of reusable services and then composes them into larger workflows.

  • Verified service construction: Coding agents turn tool specifications into service implementations. The services expose typed states and shared interfaces, making their behavior easier to connect and check.
  • World-model support: Some tools cannot be implemented or reproduced reliably. CompoWorld uses a world model for these cases, reducing dependence on fully simulated external systems.
  • Dependency-driven task generation: A random-walk procedure connects services through dependency graphs. The resulting tasks require the agent to carry information across services, preserve intermediate state, and use outputs in later actions.
  • Completion-oriented optimization: Verified trajectories are used for supervised fine-tuning. During reinforcement learning, the Completion-Focused Rubric Reward gives more attention to criteria with lower pass rates within each rollout group, encouraging the model to finish the entire workflow instead of optimizing only easy subtasks.

The authors report 448 constructed services exposing 10,130 tools. They use 3K supervised fine-tuning trajectories and 1K reinforcement-learning tasks to train Qwen3.6-35B-A3B. According to the paper summary, the resulting system improves on its backbone by an average of 9.17 points across eight benchmarks. On AutomationBench, it is reported to outperform frontier models including Claude Opus 4.6 and to lead the compared agent-specialized 35B-A3B models.

Why it matters

The main contribution is a shift in how environment scale is defined. Instead of measuring scale only by the number of tasks within one application, CompoWorld expands the space through combinations of services and dependency structures. A finite service library can therefore support many workflows, provided that its interfaces and state transitions remain verifiable.

This design is relevant to agents expected to perform practical, multi-step automation. It gives training data a stronger emphasis on information transfer, tool coordination, and end-to-end completion. It also illustrates why environment quality may matter as much as model size or instruction volume when developing general agents.

There are clear limitations to keep in view. The approach depends on the quality of generated services, the accuracy of the world model, and the coverage of automated verification. The reported benchmark gains do not by themselves establish reliability on open-web tasks, real permission systems, or high-risk business operations. More implementation detail and independent reproduction will be needed. Even so, CompoWorld offers a concrete route from isolated tool tasks toward compositional workflow training.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles