
CMI says the Millennium Prize problem has apparently been settled.
Early StationeryBench results suggest GPT-6 Astra may mark a notable advance in spatial reasoning, completing tasks with dual-arm robots that MolmoAct2 did not finish.

StationeryBench evaluates robotic spatial reasoning through desk-object tasks, including uncapping a marker, pouring out paper clips, and passing a ruler between two robot arms. The benchmark compared OpenAI’s GPT-6 Astra with Ai2’s MolmoAct2 using the same dual-arm YAM robots across 200 trials. The test setup matters because it focuses on physical task progress, not just text or image-based reasoning.
GPT-6 Astra fully completed 7 out of 100 tasks, while MolmoAct2 completed zero. Astra also reached a median progress score of 46 out of 100, compared with MolmoAct2’s 12. The results suggest Astra is substantially better at making progress in these controlled robotics tasks, though the completion rate also shows the challenge remains difficult.
Yoav Artzi, an AI researcher at Cornell and Google DeepMind, described Astra’s performance as a “step change in spatial reasoning.” The article also notes that Astra approaches human-level accuracy on the still-unpublished REMAP benchmark, while Artzi cautions that it still does not match humans in other scenarios. He suspects OpenAI may have trained the model on large amounts of 3D data, such as Blender scenes, which would align with improvements on 3D tasks.
For robotics, the key signal is not that GPT-6 Astra has solved manipulation, but that it appears to make more reliable progress on spatially grounded tasks than a direct competitor in this benchmark. The combination of videos, code, and benchmark results on GitHub gives researchers a way to inspect the claims more closely. If these early results hold up, spatial reasoning could become a more important measure of frontier AI progress beyond language performance.

CMI says the Millennium Prize problem has apparently been settled.
Top AI figures are backing a frontier AI slowdown, while critics warn about regulation, transparency, and open-source risks.

Astra beat Claude Fable 5.1 in Vending-Bench and cleared all five Drone-Bench subtasks in best attempts, though reliability remains limited.

The two-year-old robotics data startup is reportedly nearing a Sequoia-led deal.