OpenAI’s unreleased model, GPT-6 Astra, demonstrated a significant leap forward in spatial reasoning during physical robotics tests, according to evaluations on the new StationeryBench benchmark. In trials using dual-arm YAM robots across five desk-object manipulation tasks, Astra fully completed 7 out of 100 tasks and achieved a median progress score of 46 out of 100, whereas Ai2’s MolmoAct2 completed zero tasks and recorded a median score of 12.

AI researchers evaluating the results attribute Astra’s performance advantage to extensive pre-training on 3D datasets, such as synthetic Blender scenes, which significantly enhances physical spatial awareness. On the unpublished REMAP benchmark, Astra reportedly achieved spatial reasoning accuracy approaching human levels, though researchers noted performance gaps remain in complex real-world scenarios.

The benchmark findings underscore OpenAI’s expanding focus on embodied AI and physical automation as the company develops long-term plans for consumer robotics. Complete benchmark code, dataset trials, and execution videos from the evaluation have been published on GitHub.

Why it matters

  • Step-change improvements in spatial reasoning bring LLMs closer to practical deployment in physical robotics and industrial automation.

  • Training on 3D synthetic data is proving critical for bridging the gap between digital language models and physical world interaction.

Source: the-decoder.com