Back to articles
AI Agents

Inherit-MAS Evolves Multi-Agent Workflows by Reusing Structure and Execution

3 min read

Introduction

The challenge in multi-agent systems is not simply assigning several language models to one task. It is also deciding how those agents should communicate, which tools they may use, and how the overall process should improve after an unsuccessful attempt. A workflow designed in advance may fail on complex cases. Yet rewriting the entire workflow at test time can remove components that were already useful. Repeating every unchanged model or tool call creates a second problem: the system spends computation rediscovering results it already has.

An arXiv paper introduces Inherit-MAS as a way to make this trade-off explicit. Its design borrows the basic intuition of inheritance and selection: retain what appears useful, make a targeted change in response to feedback, and select better candidates over successive rounds. The method separates this idea into workflow-level inheritance and execution-level inheritance.

How the method works

  • Workflow generation and judging. A meta-model synthesizes a workflow made of worker agents. Each agent is described through its role, communication inputs, and tool permissions. After execution, a separately prompted judge scores the candidate and diagnoses its deficiencies.
  • Workflow inheritance. A normal refinement round starts from the latest completed candidate rather than generating a completely new system. The method may remove nodes judged to be unhelpful and applies a validated edit intended to address the diagnosed problem.
  • Execution inheritance. When the edited workflow runs, previously stored results are eligible for reuse only if the complete resolved request and the execution context match. This condition is designed to prevent stale or mismatched results from being transferred to a new computation.
  • Controlled change. The two mechanisms address different costs. Workflow inheritance limits unnecessary structural drift, while execution inheritance avoids repeating unchanged language-model and tool calls.

Reported results

With GPT-4o-mini workers, the authors report a 55.4% completion rate on WorkBench and a 49.7% joint F1 score on HotpotQA FullWiki. According to the paper, Inherit-MAS outperforms EvoAgent, EvoMAS, and TacoMAS on these evaluations. The authors also report that, with Qwen3-32B workers, it exceeds those evolving-MAS baselines on both benchmarks.

The efficiency comparison isolates execution inheritance by contrasting the proposed controller with the same controller run without that mechanism. Worker-token usage falls by 29.1% on WorkBench and 34.6% on HotpotQA. Total token usage falls by 5.3% and 18.1%, respectively. The gap between worker-token savings and total-token savings suggests that other components, including workflow generation and judging, still contribute substantially to the overall budget.

Why it matters

Inherit-MAS frames test-time improvement as a sequence of diagnosis, local editing, validation, and selection rather than repeated full regeneration. For agent systems that require multiple attempts, this can make refinement more stable and make expensive downstream calls more economical. The strict matching rule is particularly important: reuse is useful only when the old result remains valid for the exact resolved request and context.

The available material also leaves open several questions. The abstract does not provide detailed ablations, examples of failed edits, or a breakdown of how much each inheritance mechanism contributes independently. Practical performance may depend on cache-hit rates, judge reliability, and the cost of managing larger workflows. Still, the paper offers a clear engineering principle for evolving agent systems: modify workflows locally when possible, reuse executions conservatively, and let observable execution feedback guide the next change.

Source: arXiv

Comments

Checking sign-in status...

Loading comments...

Related articles