Claude Opus 53 mins read

Claude Opus 5 sets a new ARC-AGI-3 benchmark lead

Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, far ahead of GPT-5.6 Sol's previous 7.8 percent record, according to The Decoder's report on ARC Prize results.

Opus 5 posts a wide ARC-AGI-3 lead

Claude Opus 5 ARC-AGI-3 leaderboard chart
Image credits:ARC Prize

Anthropic's Claude Opus 5 scored 30.2 percent on ARC-AGI-3, becoming the new leader on the benchmark. The previous record was 7.8 percent, set by OpenAI's GPT-5.6 Sol (Max), while Anthropic's Fable-class models reached around 20 percent, according to ARC Prize. Opus 5 also solved five previously unsolved environments, four of them at or above human level.

Why ARC-AGI-3 matters

ARC-AGI-3 is designed to measure how well AI models handle new tasks they did not encounter during training. The current version works like a game: the model must infer the rules of an interactive environment, plan actions, and execute them step by step. Official scores count the language model's own performance, not results achieved with extra software harnesses.

Researchers point to stronger reasoning signals

The ARC Prize team attributes Opus 5's lead to stronger logical reasoning that supports more autonomous exploration, planning, and execution in unfamiliar environments. During testing, the model translated tasks into algebraic notation and independently formulated reflection equations, behavior the benchmark developers said they had not seen from another model before. The result suggests a meaningful improvement on this benchmark, especially for rule discovery and structured problem solving.

Independent tests add caution

Anthropic has not explained the gain. The Decoder notes that targeted data labeling and reinforcement learning are plausible factors because Opus 5 was developed after ARC-AGI-3 and its format became public, though that does not show it was trained on the exact tasks. Tests on Guanghan Ning's private Witness benchmark suggested narrower gains, with Opus 5 statistically tying Kimi K3 and Fable 5 and improving far less over Opus 4.8 than it did on ARC-AGI-3.

Greg Kamradt, one of the researchers behind ARC-AGI-3, said the results do not rule out broader reasoning gains. Ning later clarified that Opus 5 did show broader gains on Witness, but they were much smaller than on ARC-AGI-3. The key takeaway: Opus 5's ARC-AGI-3 jump is notable, but broader claims need more detailed results from unfamiliar tasks.

Discover More

    Claude Opus 5 series benchmark comparison
    Claude Opus 5 benchmark edge

    Opus 5 leads major AI benchmarks and can cost less than Fable 5, but reliability remains a concern.

    AnthropicClaude Opus 5
    Prompt injections and Claude illustration from The Decoder
    Opus 5 vs. Prompt Injection

    Opus 5 plus Auto Mode reportedly hit zero prompt injection success across 129 browser-agent tests.

    AnthropicClaude Opus 5
    Claude Opus 5 logo image
    Claude Opus 5 Benchmarks

    Anthropic positions Claude Opus 5 as a lower-cost flagship with strong coding and reasoning benchmark results.

    AnthropicClaude Opus 5