
LEGO-Anything shows promise, but benchmark results expose weak geometric self-assessment.
A Forecasting Research Institute interim report says leading AI experts and superforecasters were too conservative on several AI milestones, including math performance and AI company revenue, while real-world uses remain harder to predict.

The Forecasting Research Institute says leading AI experts, economists, policy specialists and superforecasters consistently underestimated several recent AI advances. Its LEAP panel included 339 experts, spanning computer scientists, industry experts, economists and AI policy specialists. The takeaway for readers: forecasts from highly qualified groups can still lag fast-moving technical progress, especially when models improve across benchmarks and adoption metrics at the same time.
The report’s sharpest miss was in math: AI reached gold-medal level at the International Mathematical Olympiad in July 2025, five years ahead of the median expert forecast and ten years ahead of the median superforecaster forecast. In virology, experts expected AI models to match a top team on a troubleshooting benchmark by 2030, while superforecasters put it at 2034; FRI says that likely happened as early as April 2025. Economic forecasts also fell short, with Anthropic’s annualized revenue described as far above the medians predicted by surveyed groups.
FRI’s findings were not uniformly about underestimation. In one controlled trial, participants using a language model and internet access completed biological lab tasks at a lower rate than those using the internet alone, and the language model made no measurable difference in that small trial. Experts may also have overshot autonomous ride-hailing, with their median forecast for autonomous US ride-hailing trips in 2027 above an LLM projection cited by FRI. That makes the practical message more nuanced: benchmark gains do not automatically translate into broad real-world deployment.
FRI plans to highlight a subgroup of respondents who expect very rapid AI progress through 2040 and publish continuously updated LLM forecasts alongside human forecasts. The institute also wants to identify the most accurate LEAP panelists as more data becomes available. FRI notes an important caveat: underestimates become visible as soon as reality passes a prediction, while overestimates are easier to confirm only after deadlines pass. Readers should treat the report as evidence that AI forecasting needs faster feedback loops, not as proof that every AI impact is arriving ahead of schedule.

LEGO-Anything shows promise, but benchmark results expose weak geometric self-assessment.

RRSI aims to help self-improving AI agents generalize beyond their test tasks while using fewer tokens.

An open-source lunar foundation model turns years of Moon observations into reusable AI tools for science.

A low-cost academic AI system defeated Stratego legend Pim Niemeijer and challenged a long-standing human edge in hidden-information board games.