AI Costs3 mins read

AI Performance Costs Are Plunging, But Model Choice Is Still More Than Price

New analysis cited by The Decoder says the cost of hitting fixed AI benchmark performance is falling rapidly, with Epoch AI estimating about 13x per year and MIT researchers putting underlying algorithmic efficiency gains at about 3x per year after controls.

The headline: fixed benchmark performance is getting much cheaper

Price per question for a fixed accuracy rate across multiple AI benchmarks over time
Image credits:Epoch AI

The Decoder reports that the price of reaching a fixed performance level on selected AI benchmarks has dropped sharply since 2023. Epoch AI estimates the decline at about 47 percent per quarter, or roughly 13x per year, while MIT researchers looking at comparable data see annual drops of 5x to 10x. The key distinction: these figures describe the cost of matching a benchmark score, not necessarily the cost of running today’s most capable model in real-world workflows.

Algorithmic efficiency is improving, but it is not the whole story

Breakdown of the 2024-2025 AI price decline into algorithmic efficiency, hardware improvements, and competition
Image credits:Gundlach et al.

MIT researchers estimate annual algorithmic efficiency progress at about 3x after stripping out cheaper hardware and competitive pricing pressure. That helps explain why Epoch AI’s 13x figure is higher: it reflects market prices for benchmark performance rather than only architecture or algorithm gains. For readers comparing models, the takeaway is simple: falling prices are real, but the reason behind the drop matters.

Reasoning models can still cost more per task

Cheaper benchmark performance does not mean every newer or stronger model is cheaper to use. The Decoder notes that reasoning models can burn through far more compute per task, which can raise cost per query even as per-token prices fall. A model may score higher because it spends more processing power, not because it is more efficient in a narrow cost sense.

How to evaluate models beyond token price

Price per token is only one part of model selection. The report points to quality, latency, context window, output speed, and error rate as practical factors that can matter as much as price. A cheap model may be a poor fit for a real-time chatbot if latency is high, while a pricier model may save money if it reduces retries by getting more answers right the first time.

Discover More

    Claude Opus 5 ARC-AGI-3 benchmark result illustration
    Opus 5 leads ARC-AGI-3

    Claude Opus 5 posts a major ARC-AGI-3 score jump, with caveats from independent tests.

    Claude Opus 5ARC-AGI-3
    Claude Opus 5 series benchmark comparison
    Claude Opus 5 benchmark edge

    Opus 5 leads major AI benchmarks and can cost less than Fable 5, but reliability remains a concern.

    AnthropicClaude Opus 5