AI Is Starting to Build AI: How the RSI Race Is Splitting
Introduction
AI helping to build the next generation of AI is moving from a theoretical idea to an engineering problem that companies can measure. Anthropic recently disclosed internal figures claiming that, as of August 2026, Claude reached a “lead” level on about 26% of its AI research and development tasks. More than 90% of the work had reached at least an AI-collaboration stage, while roughly 30,000 agents were running simultaneously on Anthropic’s internal platform.
The word “lead” needs careful interpretation. It does not mean that a quarter of Anthropic’s research is fully autonomous, nor that researchers have been replaced. Anthropic uses a six-level automation scale. AL3 describes substantial AI assistance under close human guidance. AL4 means that, after receiving a high-level objective, AI can complete most of a task end to end, while people supervise and make final decisions. Only AL5 approaches full autonomy.
Key points
- Automation is accelerating. Anthropic says the share of model-development tasks reaching AL4 rose from below 1% in February to 26% six months later. The assessment covered about 15,000 granular tasks and weighted them by human time, including pretraining, reinforcement learning, evaluation failures, and inference-service incidents.
- The measurement is useful but subjective. Claude performed the initial classification and employees conducted cross-checks. The agreement between the model and humans was higher than the agreement among different human reviewers, but the boundary between collaboration and leadership is still not completely objective.
- Agent infrastructure matters. Each internal agent receives a persistent identity, while communication takes place through a shared message system. Logs are linked to source material and execution traces, making it possible to inspect responsibility, corrections, and coordination across agents.
- The leading companies are taking different routes. OpenAI is focused on scaling automated AI researchers and increasing the amount of parallel reasoning and experimentation. Google is exploring ways for agents to improve the search process itself. Recursive emphasizes preserving useful experimental knowledge so that each research cycle can improve the next one.
Why it matters
All of these approaches point toward the same intermediate trajectory: as models improve, they can handle more coding, experiment execution, analysis, and infrastructure work. More agents can then expand the number of ideas and experiments a lab is able to test. That is not yet full recursive self-improvement.
A strict RSI loop requires a durable feedback path from model capability to model improvement. If humans still decide which questions matter, whether results are credible, and when a research direction should be abandoned, the system remains a highly automated laboratory rather than an AI that autonomously builds its successor. The central bottleneck may therefore shift from generating candidates to verifying them and preventing reward hacking or evaluation gaming.
Anthropic, OpenAI, Google, and Recursive differ in emphasis, but together they reveal a broader transformation. The AI lab is becoming a system composed of models, agents, tools, evaluators, and human judgment. The important question is not whether RSI has already arrived, but which research functions can be automated reliably and how long human judgment remains essential in the loop.
Source: InfoQ Chinese
Comments
Checking sign-in status...
Loading comments...