ASI-Bench: How Far Is AI from Conducting Research on Its Own?
Introduction
Current AI systems can search literature, summarize papers, write code, and execute well-defined workflows. Those abilities are useful, but they do not necessarily amount to scientific discovery. Real research begins where the instructions become incomplete: a researcher must identify a promising direction, choose an appropriate method, respond to unexpected results, and turn an idea into evidence that others can verify.
ASI-Bench is designed to examine this gap. The benchmark focuses on innovative exploration and autonomous scientific execution across general research domains, rather than measuring only whether a model can produce a correct answer from existing knowledge.
What the benchmark tests
- Project-level research: ASI-Bench includes 60 research projects across 11 scientific domains. The unit of evaluation is therefore broader than a single question or isolated generation.
- A guidance gradient: Systems are tested with different levels of methodological support. Guidance is progressively withdrawn until the agent must decide which method to use itself.
- Verifiable outputs: Tasks go through expert review, AI-assisted auditing, sandbox execution, and scorer validation. This is intended to distinguish executable, checkable work from text that merely sounds plausible.
- A substantial autonomy gap: Across 18 state-of-the-art agent-model configurations, the average score was 50.91 with full methodological guidance. It dropped to 29.10 when only the method was specified, and to 26.62 when the agent had to determine the method independently.
Why the decline matters
The result suggests that much of today’s AI capability remains dependent on human-provided research structure. An agent may be effective at carrying out a planned procedure while struggling to decide how a vague problem should be decomposed, which assumptions deserve testing, or how a failed attempt should change the next step.
This distinction is easy to miss in conventional evaluations. When a benchmark supplies the data, tools, procedure, and success criteria, a model mainly demonstrates knowledge retrieval and workflow execution. Scientific work is less tidy. Researchers operate with incomplete information, competing hypotheses, uncertain measurements, and methods that may need to be invented or adapted during the project.
ASI-Bench is valuable because it turns this broader challenge into a structured stress test. Its design links several capabilities that are often assessed separately: exploring an unknown problem, selecting a method, executing an investigation, and validating the result. For model developers, the benchmark can help expose where an agent breaks down. For research organizations, it is a reminder that a polished report should not be treated as evidence of reliable discovery.
The benchmark is not, by itself, a complete measure of scientific creativity. Domain differences, task design, scoring procedures, and available tools can all affect the outcome, and the abstract does not provide the full details of every task. The more defensible conclusion is narrower but important: once human methodological scaffolding is removed, current systems lose substantial capability. Building agents that can generate hypotheses, choose methods, recover from failure, and check their own claims remains a central challenge on the path toward more autonomous AI scientists.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...