Can AI Agents Do Open-Ended AI Research? Two Shadow Evaluations Offer Early Evidence
Lead
Many forecasts of rapid AI progress depend on a crucial assumption: AI agents will eventually automate AI research itself. Yet the evidence for that claim remains thin. Most existing evaluations either focus on narrow tasks with clear answers or rely on conventional peer review of AI-generated papers, a process that can be overloaded, noisy, and inconsistent.
This paper proposes a third approach: “shadow evaluations.” Instead of giving an agent a toy benchmark, researchers ask it to work on the central open-ended question behind a strong unpublished paper. The original authors of that paper then review the agent’s output, using their own deep context to judge whether it meaningfully advances the research.
Key points
- A more realistic evaluation setup: The study used two unpublished NeurIPS 2026 submissions as reference projects, asking frontier agents to address their core research questions.
- Meaningful resources were provided: The agents received six days and thousands of dollars of compute, so the exercise was not simply a low-effort prompt test.
- Engineering was not the main bottleneck: According to the paper, the agents completed the required engineering without human help.
- Research progress was limited: Despite executing the technical work, the agents did not substantially answer the underlying research questions. Both outputs were clearly rejected by the original authors.
- The pattern appeared robust: A check using a second model and scaffold reproduced the same broad failures.
The authors highlight five recurring failure modes. The agents showed weak judgment about the bar for publishable research, responded uncreatively when research designs had shortcomings, failed to backtrack effectively from dead ends, displayed poor awareness of resource constraints, and drifted away from instructions over time.
Why it matters
The paper does not claim that AI systems are useless for research. On the contrary, it suggests that today’s agents may already be able to automate a meaningful share of the engineering involved in AI R&D. The harder parts, however, are less mechanical: choosing promising directions, recognizing when an experiment is not good enough, redesigning a study, and allocating limited time and compute wisely.
For labs, policymakers, and forecasters, the implication is important. Progress toward automated AI research cannot be inferred only from coding benchmarks, closed-form tasks, or isolated demonstrations. Evaluations need to cover the full research lifecycle, especially the messy and judgment-heavy stages. Shadow evaluations are costly and still based on a small sample here, but they point toward a more grounded way to measure whether agents are becoming genuine research collaborators rather than merely capable executors.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...