
Claude Opus 5 posts a major ARC-AGI-3 score jump, with caveats from independent tests.
New analysis cited by The Decoder says the cost of hitting fixed AI benchmark performance is falling rapidly, with Epoch AI estimating about 13x per year and MIT researchers putting underlying algorithmic efficiency gains at about 3x per year after controls.


The Decoder reports that the price of reaching a fixed performance level on selected AI benchmarks has dropped sharply since 2023. Epoch AI estimates the decline at about 47 percent per quarter, or roughly 13x per year, while MIT researchers looking at comparable data see annual drops of 5x to 10x. The key distinction: these figures describe the cost of matching a benchmark score, not necessarily the cost of running today’s most capable model in real-world workflows.

MIT researchers estimate annual algorithmic efficiency progress at about 3x after stripping out cheaper hardware and competitive pricing pressure. That helps explain why Epoch AI’s 13x figure is higher: it reflects market prices for benchmark performance rather than only architecture or algorithm gains. For readers comparing models, the takeaway is simple: falling prices are real, but the reason behind the drop matters.
Cheaper benchmark performance does not mean every newer or stronger model is cheaper to use. The Decoder notes that reasoning models can burn through far more compute per task, which can raise cost per query even as per-token prices fall. A model may score higher because it spends more processing power, not because it is more efficient in a narrow cost sense.
Price per token is only one part of model selection. The report points to quality, latency, context window, output speed, and error rate as practical factors that can matter as much as price. A cheap model may be a poor fit for a real-time chatbot if latency is high, while a pricier model may save money if it reduces retries by getting more answers right the first time.

Claude Opus 5 posts a major ARC-AGI-3 score jump, with caveats from independent tests.

Opus 5 leads major AI benchmarks and can cost less than Fable 5, but reliability remains a concern.