August 19, 2026·5 min read·AIgentic.media

Why AI's Recursive Self-Improvement May Not Arrive Soon

ai-newsresearchai-researchai-agentsai-safety
Why AI's Recursive Self-Improvement May Not Arrive Soon

An abstract visualization of AI analyzing complex research data, with interconnected nodes representing failed attempts at scientific discovery

The AI industry's most seductive promise is that artificial intelligence will soon improve itself, with almost no need for human oversight. Large language models can already write code, generate synthetic training data, and optimize the computer chips they run on. Forecasts of explosive progress predict that what researchers call recursive self-improvement is on the horizon.

A new study suggests that horizon may be further away than the hype suggests.

Researchers at Princeton University, MIT, and several other institutions put Anthropic's Claude Opus 4.8 through the most rigorous test of AI research capability to date — and the AI failed. It could handle every engineering task required to conduct AI research. It could not produce original science worth publishing.

The shadow evaluation

The research team, led by Peter Kirgis and Sayash Kapoor at Princeton, designed a method they call "shadow evaluation." Instead of testing AI on narrow benchmarks with checkable answers, they asked the model to answer real research questions from two high-quality papers submitted to NeurIPS 2026 — one of machine learning's top conferences.

Because the papers had not been made public, the AI could not have memorized the answers from its training data or found them online. The agents had to produce genuinely original research.

The setup was generous. Claude Opus 4.8, running on an open-source framework called OpenClaw, received six days, $3,000 in Anthropic API credits, a GPU budget to run experiments, its own virtual computers, and unrestricted access to the web. It could spawn subagents to handle pieces of the work. It had everything a human PhD student would want.

The original authors of the two papers graded the AI's results as they would evaluate a submission to a conference. They rejected both.

Good engineers, bad researchers

The agents could do the work. They reviewed the literature, ran hundreds of experiments, and compiled results. What they could not do was think creatively about what those results meant.

"They were unambiguously bad at carrying out the research itself," says Kapoor. The agents ran bizarre experiments, in one case testing hypotheses on tiny synthetic datasets that bore no relation to real-world conditions. They struggled to write intelligibly about their work. They made no novel contribution to their fields.

The study identified a specific pattern of failure. The agents developed novel and ambitious hypotheses resembling those the original human authors started with, but then rejected them on the basis of very limited data. They committed to unpromising approaches too quickly and could not backtrack. They could make small pivots but could not fundamentally rethink their approach or start over from scratch.

Subagents occasionally hallucinated or misrepresented results, though the main orchestrator agent caught those errors. The agents did not engage in reward hacking — hiding or misrepresenting experiments to make their results look better.

What the industry says internally

The finding may echo what AI companies are finding internally, regardless of their most optimistic public statements. Anthropic cofounder Jack Clark wrote in his newsletter Import AI that the results rhyme with what the company found when it tried to automate aspects of AI safety research.

"There's a certain absence of valuable, intuitive creativity in today's AI systems," Clark wrote. "Though they're extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them from being good researchers." He called the lack of creativity a "bearish signal on short recursive self-improvement timelines."

This is not for lack of trying. OpenAI has made building an automated AI researcher an explicit goal. Anthropic identifies self-improving AI as the industry's next milestone. In June, Anthropic published a blog post titled "When AI Builds Itself," charting its progress toward models that speed up their own development. In July, OpenAI advertised that GPT-5.6 Sol had helped post-train a smaller model, saving researchers weeks of work.

The gap between the engineering and the science is stubborn.

The trillion-dollar question

The study has limitations. It covered just two research papers, and the original authors knew the papers they were grading were AI-generated, which could have colored their evaluations. The researchers had substantial discretion in designing the study, meaning their preexisting beliefs could have slipped into the results.

But the central finding — that today's most advanced AI can handle the mechanics of research but not the creativity — challenges a core assumption of the AI accelerationist camp. If the biggest advances in the field, like the invention of the transformer architecture, required genuine creative leaps, then grinding away on narrower tasks may not be enough to trigger recursive self-improvement.

"That's frankly the trillion-dollar question right now," says Kapoor.

Some researchers believe that all the ingredients for transformative AI are already present: faster training, better benchmarks, more efficient architectures. If that is true, creativity may not be necessary. The AI could improve itself through brute force optimization on tasks whose success can be automatically checked.

But if the biggest breakthroughs require the kind of open-ended thinking that today's models cannot do, then the timeline for AI improving itself may stretch much longer than the industry's most vocal optimists suggest. The tools that will build the next generation of AI may need to be smarter than the tools that build today's code.

Sources

Want to learn more?

Let's discuss how AI can transform your business.

Explore AI Agents