MILO: Evolving Agent Harnesses with Orchestrated Multi-Agent Search
Introduction
For long-horizon tasks, an AI agent is more than a language model. Its performance also depends on a harness that controls planning, tool use, execution, and interaction with the surrounding environment. A harness that works well for one model may become less effective after the model changes, forcing developers to repeat costly manual tuning. MILO, or Meta-evolutionary Island Orchestration, treats this problem as an evolving search process.
What MILO changes
- It searches complete harnesses. Many automated approaches optimize only prompts, skills, or another isolated component. MILO’s mutator agents rewrite entire harnesses, allowing changes to be evaluated as coordinated systems rather than as disconnected parts.
- It preserves a structured search history. Candidate solutions are organized into island-based lineage trees. The system records not only successful descendants but also rejected mutations, using failed attempts as negative evidence to avoid repeating unproductive directions.
- It evolves the search process itself. An orchestrator can graft lineages, create or maintain different species, reassign mutator agents, and revise the search curriculum. This makes the exploration strategy adaptive instead of permanently fixed.
- It evaluates across models and tasks. The study tests MILO on Terminal-Bench 2.1, PaperBench, and DeepSWE with both the frontier model Opus 4.8 and the open-weight model gpt-oss-120b. The discovered harnesses outperform eight advanced harnesses and six competing search methods in the reported comparisons.
Why the results matter
With Opus 4.8, MILO improves resolution over the initial harness by 12.0%, 28.3%, and 10.3% on the three benchmarks, respectively. On Terminal-Bench 2.1, it reaches 86.1±2.0%, above the official leaderboard’s reported top entry of 83.8±2.3% at the time of comparison.
The broader contribution is the idea of searching for better search behavior. As harness design becomes more modular and combinatorial, a fixed optimization recipe can become exploitative or settle too quickly on a narrow family of solutions. Island structures help preserve diversity, while lineage memory connects current decisions to both successful and failed experiments. The orchestrator can therefore change not only candidate harnesses, but also how candidates are generated and evaluated.
This does not make harness discovery a solved problem. Transfer across tasks, models, and tool environments may remain difficult, while search cost, evaluation reliability, and interpretability are still practical concerns. Even so, MILO points toward a shift in agent engineering: from manually maintaining static execution recipes to building systems that continuously discover and refine their own operating procedures.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...