
CMI says the Millennium Prize problem has apparently been settled.
OpenAI’s GPT-6 Astra performed strongly on ErdosBench despite math not being the company’s stated priority, pointing to a more specialized, “spiky” path for AI progress.

GPT-6 Astra topped ulam.ai’s ErdosBench for open math problems before Fable 5.1 later retook the lead with better overall performance. The benchmark covers 226 open math problems inspired by Erdős problems, and Astra scored 3.23 while solving 106 problems, 43 completely, and disproving 27 others.
The key reader takeaway: Astra is a serious advance for AI-assisted math research, but the benchmark is not saturated. Strong results in one demanding domain do not automatically mean a model is advancing evenly across all capabilities.
OpenAI chief scientist Jakub Pachocki said the company could make models better at mathematics research with more focus, but does not prioritize that direction because of urgency around recursive self-improvement and automated alignment research. That makes Astra’s math strength notable: according to the article, it appears to be a byproduct of other priorities rather than the result of targeted math optimization.
For readers tracking AI strategy, the important point is resource allocation. Even leading AI labs appear to face trade-offs about which capabilities to push hardest.
The article frames Astra as evidence for a jagged or “spiky” AI trajectory, where models become extremely strong in select areas such as coding or math while other abilities may stagnate or improve more slowly. This contrasts with a mainstream AGI view in which systems improve gradually and broadly across human tasks.
The practical implication is simple: benchmark wins need to be read by domain, not treated as universal capability gains. A model that dominates math may not necessarily dominate language quality, common sense, or social reasoning.
The article also highlights a broader concern raised by mathematician Terence Tao at the 2026 International Congress of Mathematicians: if AI produces proofs faster than humans can verify them, math could move from scarcity to overload. In that scenario, the central challenge becomes deciding which results matter, not just generating more results.
The hardest math problems remain unsolved for now, which gives the field time to adapt. But the direction is clear: AI is forcing mathematicians to rethink verification, value, and the role of human understanding.

CMI says the Millennium Prize problem has apparently been settled.
Top AI figures are backing a frontier AI slowdown, while critics warn about regulation, transparency, and open-source risks.

AllSpark’s Qwen-based Iris-mini and Iris-pro post leading open-weight search-agent results, according to the paper.

Astra completed 7 of 100 StationeryBench tasks, while MolmoAct2 completed none.