AI Research4 mins read

OpenAI’s GPT-5.6 Sol Ultra Reportedly Proves a 50-Year-Old Math Conjecture

OpenAI says GPT-5.6 Sol Ultra produced a proof of the Cycle Double Cover Conjecture in under an hour using 64 parallel subagents, raising new questions about AI reasoning, verification, and citation practices.

The core claim: a long-open graph theory problem, solved fast

OpenAI says GPT-5.6 Sol Ultra generated a complete proof of the Cycle Double Cover Conjecture, a graph theory problem that had remained unproven for about 50 years. The model reportedly finished in just under an hour while using 64 subagents working in parallel. The conjecture asks whether cycles can be found in a network of vertices and edges so that each edge is traversed exactly twice.

Why mathematicians are watching verification closely

Mathematician Thomas Bloom of the University of Manchester called the proof “very nice” and described it as short, elementary, and something that could have been discovered in the 1980s. He also said a full mathematical verification by the scientific community is still pending. The immediate takeaway: the result is potentially important, but the proof still needs careful review before it can be treated as settled mathematics.

The citation problem is part of the story

Bloom criticized the proof for not citing earlier work that he says likely influenced its strategy, including a 1983 paper by Bermond, Jackson, and Jaeger. He framed this as a recurring issue with AI-generated proofs and papers: models may use ideas and proof strategies from the literature without clearly attributing them. That makes the result not just a math story, but also a test case for research transparency in AI-assisted discovery.

The prompt engineered persistence, not just intelligence

The human-written prompt reportedly pushed the system toward sustained problem-solving by instructing it to assume a complete proof exists and reject partial results or summaries. It also used adversarial checking and kept most of the 64 agents unaware of which approach seemed most promising, encouraging independent attempts. This suggests the breakthrough, if verified, may reflect a combination of model capability, parallel search, strict evaluation, and carefully designed prompting.

The bigger question: recombination or new discovery?

The case sharpens a central debate around reasoning models: are they creating new mathematical insight, or recombining known work with unusual persistence? Bloom appears to lean toward the latter for this proof, while also suggesting AI may find more open problems that require established theory plus patience. For readers, the practical lesson is to separate the achievement from the unresolved questions around novelty, attribution, and peer verification.

Discover More