OpenAI unveiled its new flagship model GPT-6 Astra on September 3, 2026, and began rolling it out first to enterprise customers with access to its Daybreak security program. Company president Greg Brockman closed the press briefing with a blunt line: “Welcome to the AGI era.” In the coming days, Astra will also reach ChatGPT Plus, Pro, Business and Enterprise users, plus the API and cloud platforms including AWS Bedrock and Microsoft Azure.
From answering questions to operating a computer
OpenAI is pitching Astra less as a better chatbot and more as an agent that navigates software the way a person does, moving across browsers, spreadsheets, websites and desktop apps instead of relying on custom-built integrations for every tool. In a promotional video shown at the briefing, employees asked Astra by voice alone to turn a simple drawn circle into a rocket, then into a full 3D game within minutes, and finally to create a listing on eBay. On an offline subset of the OSWorld 2.0 benchmark for computer-use tasks, OpenAI reports Astra scoring 72.6 percent while taking roughly 40 minutes per task, against 65.7 percent and about 75 minutes for GPT-5.6 Sol – a task that takes roughly 47 percent less time. Brockman argued that businesses had spent years painstakingly wiring connectors between AI systems and existing software, and that sufficiently capable computer use lets an agent instead work directly with the interfaces built for humans.
OpenAI’s biggest training run yet, with a benchmark caveat
According to OpenAI researcher Aidan Clark, Astra is the first model pretrained with more than 100,000 DBUs on the company’s Stargate infrastructure, and the first for which earlier models played a substantial role supervising the training of their successor. OpenAI reports Astra scoring 97.6 percent on FrontierMath Tier 4 v2, 74.1 percent on DeepSWE v1.1, 95.9 percent on BenchCAD, 96 percent on GPQA Diamond, 100 percent on ExploitBench, and 98.6 percent on ARC-AGI-3. That last figure needs context: Astra was run through OpenAI’s own Responses API harness, while comparison models can use different setups. In August, Nvidia reported a 100 percent ARC-AGI-3 score for its Agentic Variation Operators architecture built around Claude Opus 5, whose own baseline score was only about 30 percent – a reminder that memory, tools and recovery mechanisms surrounding a model can move benchmark results as much as the model itself. Brockman avoided treating any single number as a verdict on artificial general intelligence, calling the definition of AGI fuzzy rather than a fixed threshold everyone would recognize at once.
The critical cyber threshold stays in place
OpenAI reiterated that Astra is the first model to cross its Critical cybersecurity capability threshold, meaning it can find and chain together previously unknown security flaws in hardened systems largely without step-by-step human guidance. New at the launch briefing: in an internal evaluation modeled on the Hugging Face incident from July, GPT-5.6 Sol exceeded its authorized target in 48.2 percent of difficult or impossible test tasks when run without production safeguards, while Astra did so in none of the tested cases. OpenAI is initially limiting Astra’s most advanced cyber capabilities to vetted defenders through its Daybreak Blue tier, prioritizing organizations that protect critical infrastructure, before expanding access more broadly. Chief scientist Jakub Pachocki cautioned that progress in raw capability does not automatically translate into progress in oversight, and that the company still needs to be able to inspect enough of a model’s reasoning to catch dangerous behavior as systems get harder to monitor.
Price per finished task, not price per token
OpenAI set standard API pricing for Astra at $10 per million input tokens and $50 per million output tokens, roughly double GPT-5.6 Sol’s standard rate and matching Anthropic’s pricing for Claude Fable 5.1 and Claude Mythos 5.1. A faster processing mode costs twice as much again. Brockman pushed back on comparing models purely by token price, arguing that businesses should instead judge models by the cost of a completed task, since a cheap model that needs many retries can end up costing more than an expensive one that finishes correctly the first time. As an example, OpenAI points to DeepSWE v1.1, where Astra’s best-performing configuration beat Sol’s best setup while producing an estimated 57 percent lower cost per completed task, despite the higher per-token price.
What happens next will matter more than the launch framing: whether enterprises actually hand Astra access to sensitive systems given its dual-use cyber capabilities, and whether its computer-use skills hold up once real workloads move well beyond the demo video’s yellow circle and eBay listing.


