ScholarCatalyst Tests Whether AI Can Find the Papers That Spark New Research
Introduction
Scientific search is often framed as a relevance problem: given a topic or question, retrieve papers that discuss similar concepts. Research practice is more demanding. When a project is still only a rough idea, the most useful paper may not be the one with the closest vocabulary. It may be an older method, an idea from a neighboring field, or a piece of work that exposes a path around a technical obstacle.
ScholarCatalyst is designed to evaluate this less visible form of expertise. Its target is the “catalyst paper”: earlier work that did help a research project, or could have advanced it if the researchers had found it at the right time.
How the benchmark was built
The researchers asked 184 lead authors to look back at 207 recent computer science projects. For each project, authors identified papers that had contributed to the work or might have been useful, and provided detailed rationales for their decisions. The resulting resource contains about 1,000 queries as well as author-labeled hard negatives—papers that may look relevant but were not judged to be useful catalysts.
The temporal setup is central. Given the initial research question, a system must search only literature that was available when the project started. This prevents the completed project from leaking the answer backward into the retrieval process and better reflects how literature discovery happens in practice.
Main findings
- Embedding retrieval records 0.48 Recall@20.
- An agentic search system scores 0.42, despite using the same retriever as a tool.
- Claude Fable 5.1 reaches 0.51 Recall@20. Even though it may have encountered the completed papers during training, it remains far from reproducing the authors’ judgments.
- Adding planning and tool use alone does not produce expert-level scientific search.
Why it matters
Most retrieval benchmarks reward topical or semantic matching. ScholarCatalyst asks a more consequential question: could this paper have changed the trajectory of the project? That distinction matters because scientific usefulness is often relational. A paper’s value depends on the problem being attempted, the project’s stage, and the bridge between an existing idea and a new research direction.
The author-centered annotations are the benchmark’s strongest feature. Project leaders can explain why a paper mattered, what gap it addressed, and why another apparently similar paper did not help. These rationales can support systems that learn more than surface relevance. They may eventually help models identify transferable mechanisms, recognize useful analogies, and justify why a recommendation deserves attention.
The benchmark is not a perfect ground truth for scientific influence. Retrospective judgments can be shaped by memory and knowledge of the final outcome, while different authors may interpret “could have helped” differently. Still, those limitations expose an important part of the problem: expert search is partly an intuitive, context-sensitive activity that conventional relevance labels rarely capture.
For future research agents, returning a ranked list will not be enough. A useful system should connect a candidate paper to the current research question, distinguish direct evidence from speculative inspiration, and operate under the information constraints faced by researchers at the beginning of a project. ScholarCatalyst offers a concrete starting point for training and evaluating that capability.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...