Open AI3 mins read

GPT-6 Astra Shows Early Spatial Reasoning Gains in Robotics Benchmark

Early StationeryBench results suggest GPT-6 Astra may mark a notable advance in spatial reasoning, completing tasks with dual-arm robots that MolmoAct2 did not finish.

Robot arms representing GPT-6 Astra robotics and spatial reasoning benchmarks
Image credits:The Decoder

What the Benchmark Tested

StationeryBench evaluates robotic spatial reasoning through desk-object tasks, including uncapping a marker, pouring out paper clips, and passing a ruler between two robot arms. The benchmark compared OpenAI’s GPT-6 Astra with Ai2’s MolmoAct2 using the same dual-arm YAM robots across 200 trials. The test setup matters because it focuses on physical task progress, not just text or image-based reasoning.

Astra’s Early Lead Over MolmoAct2

GPT-6 Astra fully completed 7 out of 100 tasks, while MolmoAct2 completed zero. Astra also reached a median progress score of 46 out of 100, compared with MolmoAct2’s 12. The results suggest Astra is substantially better at making progress in these controlled robotics tasks, though the completion rate also shows the challenge remains difficult.

Why Researchers Are Paying Attention

Yoav Artzi, an AI researcher at Cornell and Google DeepMind, described Astra’s performance as a “step change in spatial reasoning.” The article also notes that Astra approaches human-level accuracy on the still-unpublished REMAP benchmark, while Artzi cautions that it still does not match humans in other scenarios. He suspects OpenAI may have trained the model on large amounts of 3D data, such as Blender scenes, which would align with improvements on 3D tasks.

The Practical Takeaway

For robotics, the key signal is not that GPT-6 Astra has solved manipulation, but that it appears to make more reliable progress on spatially grounded tasks than a direct competitor in this benchmark. The combination of videos, code, and benchmark results on GitHub gives researchers a way to inspect the claims more closely. If these early results hold up, spatial reasoning could become a more important measure of frontier AI progress beyond language performance.

Discover More

    OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei
    AI Slowdown Debate

    Top AI figures are backing a frontier AI slowdown, while critics warn about regulation, transparency, and open-source risks.

    AIAnthropic
    GPT-6 Astra and autonomous drone control illustration
    GPT-6 Astra tops agent tests

    Astra beat Claude Fable 5.1 in Vending-Bench and cleared all five Drone-Bench subtasks in best attempts, though reliability remains limited.

    AI agentsBenchmarks