Back to articles
AI Agents

AREX-2: Building Agents That Improve Across Long Iterative Runs

3 min read

From one-shot answers to iterative improvement

Many language-model agents can produce a plausible answer, but that does not necessarily mean they can inspect the result, learn from feedback, and improve it repeatedly. AREX-2 focuses on this missing capability. The paper defines self-improvement as the ability to refine a solution at test time through multiple iterations, rather than treating the first generated answer as final.

The authors divide this capability into two complementary parts:

  • Reflection: producing a revision that is better than the current solution.
  • Long-horizon execution: keeping the improvement process effective across many rounds, while carrying useful information from one attempt to the next.

The distinction matters. Reflection without sustained execution may result in good suggestions that are never fully implemented. Execution without meaningful reflection can turn a long trajectory into repeated, low-value actions. AREX-2 hypothesizes that both capabilities are sufficiently domain-agnostic to be learned in settings that provide clearer supervision than open-ended real-world work.

Learning from tasks with verifiable feedback

To create training data, the researchers synthesize long-horizon improvement trajectories from machine-learning and algorithmic-programming tasks. These domains are useful because their outcomes can be evaluated with relatively concrete feedback. A trajectory can therefore capture more than a final answer: it can record attempts, execution, evaluation, error discovery, and subsequent revision.

The resulting agent is built on Qwen3.8-27B. It reaches 81.8 on MLE-bench Lite and 70.7 on Frontier-CS. The reported transfer results are particularly notable: the system scores 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA. These figures suggest that behavior learned from technical tasks can extend to deep research, at least on the reported benchmarks. The abstract also states that performance keeps improving as the available number of rounds increases, indicating that the agent can make productive use of additional iteration budget.

Why it matters

AREX-2 points to a training route that is different from simply making a model larger. Instead of optimizing only for a strong first response, it treats the trajectory of problem solving as a central learning target. This could be valuable for work that naturally requires experiments, verification, debugging, and successive revisions.

The available material also leaves important questions open. It does not provide the full recipe for synthesizing trajectories, the detailed baselines, the cost of long-horizon inference, or the per-round improvement curves. The paper and accompanying code will therefore be needed to determine how much of the gain comes from reflection data, sustained execution training, or their combination.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
Why LLM Agents Keep Following Instructions Users Have Withdrawn
AI Agents
cctest.ai
AI Agents

Why LLM Agents Keep Following Instructions Users Have Withdrawn

A new study turns intent drift into a measurable failure mode: superseded or withdrawn user requirements can continue to shape an agent’s answer or tool action. Its IntentFlux benchmark and StateForge method show that explicit state maintenance helps, but does not eliminate the gap between evolving dialogues and direct single-turn tasks.

Read more