Back to articles
Coding AI

WideSWE Tests Whether Coding Agents Can Coordinate Across Repositories

3 min read

Introduction

Coding-agent evaluation is moving beyond isolated issue resolution toward longer development tasks. Yet many existing benchmarks still assume that the entire problem lives inside one repository. That assumption does not match how software ecosystems are built. A feature or bug fix may require synchronized changes in a core library, a client, an integration project, or a test repository. Editing only one of them can produce a locally plausible patch that fails at the ecosystem level.

WideSWE is designed to measure this broader capability. The authors mined and reviewed changes from 103 software ecosystems, producing 120 real-world tasks: 60 bug fixes and 60 feature requests. Prompts were derived from related issues and pull requests. The benchmark developers also reviewed and adapted hidden tests so that multiple valid implementations could pass while required behavior and regression checks remained enforced.

Key findings

  • The unit of evaluation is an ecosystem, not a single repository. Agents must identify every project affected by a request and coordinate compatible changes across them.
  • The task mix reflects ordinary engineering work. The benchmark includes an equal number of fixes and features, all based on changes observed in real software development.
  • End-to-end completion remains difficult. Across seven agent configurations, full-task success ranged from 10.83% to 42.50%. The strongest reported configuration paired Codex CLI with GPT-5.6-sol.
  • Failures involve planning as well as implementation. Agents sometimes failed to discover a required repository, recognized the work but left it incomplete, or modified the relevant repositories without fully meeting the requested behavior.
  • Execution strategy matters. Independent, one-repository-at-a-time execution mainly helped recover omitted work. It was less effective at repairing implementations that had already been attempted unsuccessfully. Joint execution could use information from related repositories to guide implementation and verification.

Why it matters

WideSWE highlights the gap between being a capable local code editor and functioning as a software-ecosystem engineer. Cross-repository work requires a global change plan, dependency awareness, synchronized interfaces, and verification that spans project boundaries. The difficulty is therefore not simply the sum of several single-repository tasks; the agent must understand how the repositories fit together and how a change in one affects the others.

The benchmark also raises an important design question for agent workflows. Isolating repositories can make execution easier to control, but it may remove context that is essential for planning and debugging. Joint execution, shared observations, and persistent task state may provide better support for coordinated work, although the results do not suggest that joint execution solves the problem by itself.

For future evaluations, WideSWE offers a more realistic target: judge whether a collection of related changes jointly satisfies the original request, rather than checking only whether one repository’s tests pass. Its results turn common but hard-to-measure problems—omissions, unfinished edits, and incomplete behavior—into explicit evaluation dimensions.

Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
From One to Many and Back: Category-Aware Training for SWE Agents
Coding AI
cctest.ai
Coding AI

From One to Many and Back: Category-Aware Training for SWE Agents

A new study addresses the uneven progress that can emerge when heterogeneous software-engineering tasks are trained in a shared reinforcement-learning pool. It develops category experts with an iterative Refresh–Repair–Expand loop, then consolidates them into one deployable model through label-routed multi-teacher on-policy distillation.

Read more
CCTest · Blog
CodeMidas Turns Source Code into Reinforcement Learning Environments for Coding Agents
Coding AI
cctest.ai
Coding AI

CodeMidas Turns Source Code into Reinforcement Learning Environments for Coding Agents

Xiaomi MiMo’s CodeMidas pipeline uses existing source code to discover functionality, generate executable tests, and filter coding tasks for reinforcement learning. The resulting tasks improved an agent’s performance across software repair, whole-program construction, and terminal benchmarks.

Read more