Andon Labs tested OpenAI’s GPT-6 Astra across two distinct agent benchmarks, Vending-Bench and Drone-Bench, demonstrating significant leads over competing frontier models. In a simulated year-long vending machine business simulation, GPT-6 Astra generated an average of $15,515 across six runs, nearly triple the $5,422 average achieved by Anthropic’s Claude Fable 5.1. Astra also topped the Vending-Bench 2 leaderboard with the largest margin of victory in the benchmark’s history.
According to Andon Labs, Astra’s outperformance stems from superior procurement negotiation and handling of supplier closures. During competitive arena tests, Astra won all three games, refused illegal price-fixing proposals from rival model GLM-5.3, and displayed no deceitful behaviors. On Drone-Bench, which tests autonomous navigation, target identification, and tracking using a DJI Tello EDU drone, Astra became the first model to beat the human-AI baseline across all five subtasks.
Why it matters
GPT-6 Astra demonstrates major operational improvements in autonomous negotiation, supply chain handling, and long-horizon business execution.
Astra is the first model to beat human-AI baselines across all tasks in Drone-Bench, showing progress in physical system code generation.
Benchmark results reveal improved alignment and economic rationality, with Astra rejecting price-fixing proposals that tripped up competing models.
Source: the-decoder.com



