
User demos show Opus 5 creating playable 3D prototypes with code-generated assets, physics, and music.
Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, far ahead of GPT-5.6 Sol's previous 7.8 percent record, according to The Decoder's report on ARC Prize results.


Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, becoming the new leader on the benchmark. The previous record was 7.8 percent, set by OpenAI's GPT-5.6 Sol (Max), while Anthropic's Fable-class models reached around 20 percent, according to ARC Prize. Opus 5 also solved five previously unsolved environments, four of them at or above human level.
ARC-AGI-3 is designed to measure how well AI models handle new tasks they did not encounter during training. The current version works like a game: the model must infer the rules of an interactive environment, plan actions, and execute them step by step. Official scores count the language model's own performance, not results achieved with extra software harnesses.
The ARC Prize team attributes Opus 5's lead to stronger logical reasoning that supports more autonomous exploration, planning, and execution in unfamiliar environments. During testing, the model translated tasks into algebraic notation and independently formulated reflection equations, behavior the benchmark developers said they had not seen from another model before. The result suggests a meaningful improvement on this benchmark, especially for rule discovery and structured problem solving.
Anthropic has not explained the gain. The Decoder notes that targeted data labeling and reinforcement learning are plausible factors because Opus 5 was developed after ARC-AGI-3 and its format became public, though that does not show it was trained on the exact tasks. Tests on Guanghan Ning's private Witness benchmark suggested narrower gains, with Opus 5 statistically tying Kimi K3 and Fable 5 and improving far less over Opus 4.8 than it did on ARC-AGI-3.
Greg Kamradt, one of the researchers behind ARC-AGI-3, said the results do not rule out broader reasoning gains. Ning later clarified that Opus 5 did show broader gains on Witness, but they were much smaller than on ARC-AGI-3. The key takeaway: Opus 5's ARC-AGI-3 jump is notable, but broader claims need more detailed results from unfamiliar tasks.

User demos show Opus 5 creating playable 3D prototypes with code-generated assets, physics, and music.

Opus 5 leads major AI benchmarks and can cost less than Fable 5, but reliability remains a concern.

Opus 5 plus Auto Mode reportedly hit zero prompt injection success across 129 browser-agent tests.

Anthropic positions Claude Opus 5 as a lower-cost flagship with strong coding and reasoning benchmark results.