Astra Solved What Mathematicians Couldn't

The math problems that stumped everyone
For at least a decade, mathematicians stared at ten open problems in fields ranging from high-dimensional geometry to quantum complexity. Nobody cracked them. Then an AI system that OpenAI calls Astra generated proofs for all of them in a single run, at a compute cost of roughly $2,000.
Noam Brown, one of the researchers behind the test-time reasoning technology used by Astra, announced the results on August 1. "An internal version of Astra, OpenAI's next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science," he wrote. "We believe it will be a major step for scientific reasoning." The post accumulated 8.4 million views.
OpenAI confirmed the results in a blog post titled "Ten advances in mathematics and theoretical computer science," which also revealed the Astra name for the first time. The company said an internal version of its next major model family produced arguments for all ten problems, then worked with human researchers to turn them into formal research papers.
The $2,000 proofs
The problems Astra solved cover remarkably diverse territory. They span high-dimensional geometry, coding theory, group theory, quantum complexity, lattice cryptography, and extremal combinatorics. One proof establishes the existence of non-sofic groups, resolving a major open question in group theory that had resisted progress for years.
OpenAI says the tokens used to generate all ten solutions would have cost about $2,000 at Sol's API rates. After the model produced its arguments, humans worked with the same model to turn them into research papers. The model also formalized each proof in Lean, the interactive theorem prover, creating machine-checkable certificates of mathematical correctness. OpenAI published a walkthrough of the model's reasoning process for each solution.
The company argued that claiming human authorship for proofs generated entirely by AI would misrepresent both the system's contribution and the nature of genuine human intellectual work. It pointed to the Leiden Declaration on AI and Mathematics as a reference for how credit should be assigned in AI-assisted research.
Mathematicians react: "big news" with caveats
Thomas Bloom, a mathematician at the University of Manchester who runs erdosproblems.com, called the results "big news" on X. He considers them more significant than the counterexample to the unit distance conjecture published in May. "Maybe not bigger than a proof of unit distance would have been, but in terms of constructions, this is big," Bloom wrote.
But Bloom also rejected the idea that AI is replacing mathematicians. He argued that the claim makes little sense when the AI draws on more than a century of mathematical theory, was built by mathematicians, and was trained on everything mathematicians have ever written.
Noam Brown added a dose of perspective. "Sadly, no Millennium Prize Problems (yet)," he wrote. The Clay Mathematics Institute offers $1 million for solving each of the seven Millennium Prize Problems, but only one has been solved since the prizes were announced in 2000. Brown also noted the cost efficiency: "We didn't spend a lot on each problem. It's possible to push test-time compute much further."
The skeptic's case
Gary Marcus, a cognitive scientist and frequent AI critic, published a detailed response arguing that the excitement around Astra rests on a "fallacy of composition" -- the assumption that because a system is great at math, it is great at everything.
"Math lends itself to verification and massive amounts of cheaply produced synthetic data where you can guarantee that the answers are correct," Marcus wrote. "The same applies to coding, but it is not true in general. You can generate as many math facts as you want; you can't simulate the open-ended world."
Marcus highlighted several missing pieces of information. OpenAI has not disclosed how many conjectures were attempted before hitting ten successes. The $2,000 compute cost does not include the salaries of the human researchers who helped turn the proofs into papers. Columbia University professor Henry Yuen noted that the proof writeups had the "characteristic" hallmarks of ChatGPT-generated text, elaborating on boilerplate setup while glossing over the hardest steps.
Ernie Davis, a computer science professor at NYU, added that the problem of autoformalization -- turning mathematics written by a human into a strictly logical form -- is not close to being solved. He pointed to the ongoing project to formalize Andrew Wiles' proof of Fermat's Last Theorem in Lean, which has taken years and is still not complete.
What Astra actually is
Astra is not a single model but a new model family designed to work on problems for hours or days by coordinating multiple AI agents. CEO Sam Altman demoed the system to politicians and regulators in Washington, D.C., the same week the math results were published.
The Astra family would form a new model class alongside OpenAI's existing Sol, Terra, and Luna families. Whether it ships as GPT-6 or as a variant within the GPT-5 line has not been decided. There is no release date.
Astra would be the first model tested under the Trump administration's planned new AI framework, which would require AI models to be submitted to the federal government before public release. The administration aims to finalize the framework by the end of this week.
OpenAI's long-term goal is fully autonomous AI research. Chief Scientist Jakub Pachocki said on the company's podcast that OpenAI wants to build systems that can work on a problem for hours or days. By March 2028, the company aims to have a fully autonomous AI researcher that can run research projects on its own. As early as September, OpenAI plans to have an AI system with research-intern-level skills.
Grounded, not magical
The Astra results are genuinely impressive. Solving ten open problems that had stalled for a decade, at a compute cost of $2,000, with formal verification in Lean, is a milestone worth paying attention to.
But the leap from "AI can solve hard math problems" to "AI is a universal solver" is a leap the evidence does not support. Math is a special case. It has clear rules, verifiable answers, and an infinite supply of training data with guaranteed correctness. The real world does not work that way.
The actual test of Astra will come when it is deployed outside the controlled conditions of mathematical proof, doing things like reading PDFs reliably, writing coherent scripts, or navigating open-ended tasks without compounding errors over hours of runtime. Those are the problems that still defeat every AI system, including this one.
For now, the message is straightforward: a machine produced novel mathematical proofs that humans had been stuck on for years, and it cost about as much as a decent laptop. That is a big deal. It is not the Singularity.
Sources
- The Decoder: OpenAI announces its "next major model" Astra by dropping ten previously unsolved math solutions
- OpenAI: Ten advances in mathematics and theoretical computer science
- Gary Marcus: OpenAI's amazing -- but vastly oversold -- new model Astra
- BleepingComputer: OpenAI teases Astra, its next major AI model, after it solves 10 long-standing math problems
- Gizmodo: OpenAI Smuggled the Announcement of Astra, Its Next AI Model, Into a Blog Post About Math