The independent testing laboratory Andon Labs has for the first time specifically measured OpenAI’s model GPT-6 Astra in drone control and simulated trading. In five drone sub-tasks, Astra is the first model to exceed the benchmark, but the complete chain succeeds error-free in only 2.8 percent of all attempts. In simulated goods trading, the model achieves almost three times the final capital of Anthropic’s Claude Fable 5.1.
Astra navigates offices and tracks individuals
The test Drone-Bench by Andon Labs examines whether a language model can write code for an autonomous surveillance drone solely from camera images. Five sub-tasks form the chain: creating a 3-D model of the office with obstacles from video material, determining its own position in space, navigating collision-free between rooms, recognizing a specific person based on a reference photo, and subsequently tracking them in the image.
GPT-6 Astra, which OpenAI president Greg Brockman referred to in September as the beginning of an “AGI era,” surpasses the respective benchmark in all five individual tasks for the first time, according to Andon Labs. In person recognition, this is successful in four out of ten runs, while in 3-D reconstruction it is only successful in one out of ten. When all five steps are chained into a single run, the success probability drops to 2.8 percent – just one error breaks the entire chain. From the progress of the past two years, Andon Labs derives the forecast that a top model could solve all five tasks in a single attempt by the first quarter of 2027.
Trading test shows significant lead over Fable
In the second test, Vending-Bench 2, the model manages a virtual vending machine over a simulated business year: starting capital of 500 dollars, plus purchasing, inventory management, and pricing. Over six runs, Astra achieves an average final capital of 15,515 dollars, while Claude Fable 5.1 only reaches 5,422 dollars. Even the weakest Astra run with 13,272 dollars exceeds the best Fable run with 9,874 dollars.
Andon Labs explains the lead primarily with two effects: Astra negotiates supplier prices more consistently, saving 6,152 dollars on goods purchases, while Fable repeatedly incurs losses from unreliable suppliers despite better knowledge – a total of 14,331 dollars. Additionally, Astra avoids 2,389 dollars in losses from failed supplier payments. In a separate arena variant with multiple competing AI agents, Astra wins all three test runs and independently rejects price collusion with the simulated competitors. Andon Labs, which has already tested the decisions of AI systems in the role of a leader with similar test series, sees this as evidence of how differently reliable models act in multi-stage economic decisions.
It remains open who retains control when a model increasingly reliably controls drones and makes trading decisions without a human checking every step. OpenAI had already classified Astra as a critical cyber risk before the official launch and restricted access to vetted partners in the Daybreak program – the new benchmark results now provide concrete numbers on the operational capability outside of pure text tasks for the first time. The real crux is whether oversight mechanisms can keep pace with the speed predicted by Andon Labs before a model reliably masters the entire task chain.


