Security

GPT-6 Astra flies surveillance drones – but rarely without errors

3 min read

TL;DR Too Long; Didn’t read

Andon Labs publishes two independent tests on OpenAI's model GPT-6 Astra on September 13, 2026. As a drone pilot, it passes all five test steps individually for the first time, but it only succeeds completely error-free once in 36 attempts. At the same time, Astra nearly triples the profit of Anthropic's Claude Fable 5.1 in simulated merchandise trading.

A drone with an OpenAI logo tracks a person in an office hallway while holding dollar bills in a claw. Image generated with GPT Image 2

Key takeaways

  • Andon Labs tests GPT-6 Astra for the first time in drone control and simulated trading.
  • Astra surpasses the benchmark value from Andon Labs in all five drone sub-tasks for the first time.
  • The complete task chain is only successfully completed error-free in 2.8 percent of all attempts.
  • In the trading test, Astra achieves $15,515, while Claude Fable 5.1 only achieves $5,422.
  • Astra saves $6,152 through better supplier negotiations and rejects price collusion.
  • Andon Labs expects fully solved drone tasks no earlier than the first quarter of 2027.

The independent testing laboratory Andon Labs has for the first time specifically measured OpenAI’s model GPT-6 Astra in drone control and simulated trading. In five drone sub-tasks, Astra is the first model to exceed the benchmark, but the complete chain succeeds error-free in only 2.8 percent of all attempts. In simulated goods trading, the model achieves almost three times the final capital of Anthropic’s Claude Fable 5.1.

Astra navigates offices and tracks individuals

The test Drone-Bench by Andon Labs examines whether a language model can write code for an autonomous surveillance drone solely from camera images. Five sub-tasks form the chain: creating a 3-D model of the office with obstacles from video material, determining its own position in space, navigating collision-free between rooms, recognizing a specific person based on a reference photo, and subsequently tracking them in the image.

GPT-6 Astra, which OpenAI president Greg Brockman referred to in September as the beginning of an “AGI era,” surpasses the respective benchmark in all five individual tasks for the first time, according to Andon Labs. In person recognition, this is successful in four out of ten runs, while in 3-D reconstruction it is only successful in one out of ten. When all five steps are chained into a single run, the success probability drops to 2.8 percent – just one error breaks the entire chain. From the progress of the past two years, Andon Labs derives the forecast that a top model could solve all five tasks in a single attempt by the first quarter of 2027.

Trading test shows significant lead over Fable

In the second test, Vending-Bench 2, the model manages a virtual vending machine over a simulated business year: starting capital of 500 dollars, plus purchasing, inventory management, and pricing. Over six runs, Astra achieves an average final capital of 15,515 dollars, while Claude Fable 5.1 only reaches 5,422 dollars. Even the weakest Astra run with 13,272 dollars exceeds the best Fable run with 9,874 dollars.

Andon Labs explains the lead primarily with two effects: Astra negotiates supplier prices more consistently, saving 6,152 dollars on goods purchases, while Fable repeatedly incurs losses from unreliable suppliers despite better knowledge – a total of 14,331 dollars. Additionally, Astra avoids 2,389 dollars in losses from failed supplier payments. In a separate arena variant with multiple competing AI agents, Astra wins all three test runs and independently rejects price collusion with the simulated competitors. Andon Labs, which has already tested the decisions of AI systems in the role of a leader with similar test series, sees this as evidence of how differently reliable models act in multi-stage economic decisions.

It remains open who retains control when a model increasingly reliably controls drones and makes trading decisions without a human checking every step. OpenAI had already classified Astra as a critical cyber risk before the official launch and restricted access to vetted partners in the Daybreak program – the new benchmark results now provide concrete numbers on the operational capability outside of pure text tasks for the first time. The real crux is whether oversight mechanisms can keep pace with the speed predicted by Andon Labs before a model reliably masters the entire task chain.

Frequently asked questions

Who tested GPT-6 Astra in these tests?

The independent testing laboratory Andon Labs developed the benchmarks and conducted the evaluation without the involvement of OpenAI.

Is GPT-6 Astra available to all users?

No, OpenAI continues to restrict Astra's advanced capabilities to vetted partners in the Daybreak program, a broader launch date is still pending.

Were real flying devices used in the drone tests?

According to Andon Labs, the tasks are based on video and control data from real drone flights in office environments, but each sub-task is evaluated individually with predetermined interim results, not as a continuous live flight.

How do other AI models compare?

Andon Labs primarily compares Astra with Anthropic's Claude Fable 5.1 and describes Astra as the first OpenAI model in first place on Vending-Bench 2.

Are the published figures independently verified?

The figures come directly from Andon Labs, the developer of both benchmarks; an examination by independent third parties has not yet been conducted.

Sources (2)
  1. Andon Labs: Drone-Bench Evaluation
  2. Andon Labs: GPT-6 Astra on Vending-Bench 2

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog