Can Agent Optimization Compound? A Continual-Learning Test on Terminal-Bench 2.0
A new arXiv paper argues that one-shot benchmark gains are not enough to judge agent optimizers. In a two-phase continual-learning setup, methods diverged sharply once new tasks were introduced.
Read more