EvoDuet Makes Web Search and Problem Solving Co-Evolve
Introduction
Large language models can be used as evolutionary search engines: generate a candidate, evaluate it, and use the resulting feedback to produce a better candidate. This approach becomes difficult when progress depends on knowledge outside the model’s internal repertoire. Adding a web-search tool may provide missing information, but a static retrieval routine can repeatedly surface the same pages even as the candidate solutions change.
EvoDuet addresses this mismatch by making search and solution generation co-evolve. The method keeps model parameters fixed and instead updates the interaction between candidates, queries, retrieved documents, and evaluation outcomes.
How the framework works
- Retrieval gating: At each iteration, the model assesses whether it has a knowledge gap. It can request new documents, reuse material already stored, or continue without retrieval. This makes search conditional rather than automatic.
- An inner search loop: Queries are refined and documents are ranked according to the solution scores they are expected to support. Relevance is therefore tied to downstream solving value, not only to textual similarity.
- An outer solving loop: Candidates are generated in parallel from selected documents. Their actual task scores are then recorded and fed back into later search decisions.
- Scaffold portability: The authors also report gains when EvoDuet is combined with other evolutionary-search scaffolds, including Top-K and EvoX, on Sums/Diffs and Denoising.
Results and limitations
Across 21 optimization tasks with one candidate produced per iteration, EvoDuet increased OpenEvolve’s normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna. With Gemini-3.8-Flash, the gain rose from 61.3% to 82.3%. The paper reports that its best runs exceeded previously reported best scores on eight tasks and matched them on three others, including Swap Reduction on Q20 and Rosetta.
The result is not universal. Qwen3.5-9B did not benefit from EvoDuet, suggesting that effective retrieval planning requires more than access to a search interface. The model must recognize when information is missing, formulate useful queries, interpret documents, and turn them into higher-scoring candidates.
Why it matters
EvoDuet’s main contribution is to move retrieval feedback inside the optimization loop. Search history is no longer merely a cache of context: document usefulness is connected to measured solution quality, and that connection influences future queries. This is particularly relevant for scientific optimization, where a retrieved fact or method matters only if it leads to a better experimentally scored solution.
The approach also exposes practical questions. Repeated retrieval can increase cost, poor documents may contaminate later searches, and the reported model-dependent gains suggest that search orchestration cannot fully compensate for weak reasoning. Further work should test robustness, retrieval efficiency, and performance on broader scientific problems.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...