OpenAI officially classifies its yet-to-be-released model Astra as the first system with a critical cyber capability – the highest level in its in-house Preparedness Framework. In the ExploitBench testing procedure, Astra solves all tasks and independently exploits two previously unknown software vulnerabilities. Advanced hacking features are initially available only to vetted partners in the Daybreak security program.
Astra builds complete attack chains in testing
For the current classification, OpenAI relies for the first time on concrete technical evidence rather than just a cautious risk assessment. In early August, the company could not rule out a potential critical cyber capability in Astra and paused parts of its development. In the ExploitBench test, which measures the ability to develop software exploits, Astra now masters all tasks, according to OpenAI. Against roughly twenty hardened, high-risk vulnerabilities, the model also independently turns two previously unknown zero-day flaws into a working attack chain (independently unverified). OpenAI says it has already notified the affected software vendors.
In further tests, Astra built a complete browser breakout chain: when a prepared system opened an HTML file, the model escaped the browser sandbox and executed commands on the host system. In a hardened operating system, Astra also chained several smaller flaws into a privilege escalation that grants root access from a regular user account – both without human guidance at the individual steps. For the evaluation, OpenAI says it brought in external security experts, who replicated the attack chains under controlled conditions.
Access to the strongest features stays limited
Ahead of a broad release, OpenAI is pursuing a phased approach. A small group of selected testers gets access first, followed by inclusion in the existing Daybreak security program, whose base tier Blue currently gives external defenders access to models with relaxed safeguards. Astra’s advanced cyber capabilities remain, according to Security Boulevard, reserved exclusively for vetted partners in this coalition. OpenAI has not given a date for a broadly accessible version, nor pricing or a separate arrangement for Germany and the EU. The release is reportedly set to follow soon.
In parallel, the company is tightening its technical safeguards. A monitor continuously tracks the model’s chain of thought and is meant to automatically detect and halt unauthorized actions. Faced with malicious prompts, Astra reportedly refuses to answer 91.5 percent of the time, compared with 59 percent for the predecessor model GPT-5.6 Sol. Accounts with a higher risk profile receive additionally restricted responses, backed by tests against possible breakouts during the model’s own training.
Criticism of the self-assessment grows
Astra itself was not involved in the Hugging Face breach in July, OpenAI stresses – but the lessons from that incident have fed into the new safeguards. How solid these precautions really are remains disputed. Yona Shavit, formerly at OpenAI and now with the OpenAI Foundation, voiced doubts about Astra’s rule compliance, according to TechCrunch. It could reflect genuine safety – or simply the evaluators’ expectations. An external confirmation of the classification is still pending; as in August, it rests solely on OpenAI’s own assessments.
The classification fits into a series of similar warnings. Just at the end of August, OpenAI, Anthropic, and more than a hundred other companies warned in an open letter of a window of only a few more months to prepare for cyber defense. The letter called for higher investment, shared threat data, and stronger government support for defenders of critical infrastructure. Astra now provides the most concrete, publicly documented example yet of the capability growth described there.
It remains open whether the phased, partner-limited access actually prevents credentials for Astra’s most powerful features from falling into the wrong hands. A model that builds attack chains this reliably would, after all, be valuable to attackers too. It is also unclear whether and when security agencies or companies from Germany and the EU will be admitted to the Daybreak partner circle, which OpenAI has not fully disclosed publicly.


