In Mathematical Discovery, Finding the Right Problem May Be the Real Bottleneck
Introduction: the bottleneck comes before the proof
As AI systems become more capable at research-level mathematics, the limiting factor is not only model reasoning. Expert attention is even harder to scale: mathematicians must determine whether a problem is correctly stated, whether it is genuinely open, and whether a proposed argument survives careful scrutiny. In many existing workflows, people choose a problem first, let a model work on it, and then inspect the output. Both ends of that process can become overloaded.
The paper The Problem Is the Problem: Towards Scalable Mathematical Discovery proposes a different arrangement. Instead of asking experts to name one problem in advance, it asks them to provide a research direction that matches their interests and expertise. The system then searches the literature for candidate problems within that direction.
The core idea: a staged literature-to-review cascade
The pipeline is called FAR, after its three stages:
- Find extracts conjectures, open questions, and surrounding context from a broad mathematical corpus.
- Attempt applies model reasoning to candidates that have passed initial checks, seeking possible proofs, resolutions, or useful progress.
- Recommend uses automated triage to select artifacts that deserve expert inspection.
The combinatorics pilot gives the workflow a concrete scale. It began with 5,245 papers and recovered 6,453 candidate conjectures or open problems. After filtering for apparently clear formulations and continued open status, 4,717 candidates remained. Later reasoning and triage stages surfaced 598 potential resolutions and selected 77 for author-team review. The reported examples include work related to conjectures and questions associated with Davies–Jenssen–Perkins–Roberts, Erdős–Straus, Ikenmeyer–Pak–Panova, and Lund–Saraf–Wolf.
Why it matters—and what it does not prove
FAR changes the role of AI from answering a problem chosen by a human to helping decide which problems deserve an attempt in the first place. That shift is important because research discovery contains a large amount of search, deduplication, status checking, and low-value triage. Automating those steps can reserve expensive model reasoning and expert review for a smaller, more promising set of artifacts.
The numbers should not be read as proof that the system solved hundreds of problems. “Apparently open” is not the same as definitively open, and a potential resolution still requires mathematical verification. The authors’ results therefore support FAR as a research-navigation and prioritization layer, not as an unsupervised theorem factory. Its broader lesson is that scaling AI-assisted mathematics may depend as much on managing attention and problem selection as on extending the length of model reasoning.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...