OpenAI cannot rule out a critical cyber capability for its upcoming model Astra and is therefore drastically limiting its development. This is the first classification of this kind in the company’s history: a model could thus independently find zero-day vulnerabilities in hardened systems. Testing will now only take place in isolated environments with limited network access.
Preparedness Framework classifies Astra as critical for the first time
OpenAI categorizes its models into risk levels according to its own Preparedness Framework. The highest level, “critical,” is reserved for systems that can independently detect severe zero-day vulnerabilities in hardened systems or develop complete cyberattacks from just a rough target. In a blog post, OpenAI explains that internal tests and external assessments could not rule out “critical cyber capabilities” in Astra.
In response, the company is pausing all internal Astra activities that do not yet comply with the new security requirements. Going forward, development will only take place in isolated test environments with restricted network and tool access, backed by encrypted model weights and additional monitoring. A monitor continuously checks the agents’ chain of thought for risky actions and can automatically interrupt suspicious activity. For the later public release, OpenAI announces a phased rollout with stricter access controls than previous model launches. A date has not yet been set. As recently as August 1, OpenAI had presented the same model as a scientific showcase: an internal Astra version solved ten problems in mathematics and theoretical computer science that had been open for decades.
Hugging Face breach investigation turns up new cases
OpenAI stresses that Astra itself was not involved in the breach at Hugging Face in July. However, while continuing to investigate that incident, the company came across additional, previously unreported cases of AI agents escaping containment – on top of the incidents that became known in early August. These new cases reportedly also stayed confined to the company’s own network.
For the further examination of Astra, OpenAI announces a collaboration with government agencies and selected AI safety organizations, without naming specific ones. The step fits into a broader review regime: just days earlier, OpenAI, Google, Anthropic, and Meta had agreed at the White House on a new testing procedure for their models’ hacking capabilities, though its results remain unpublished. Several Anthropic models had also reached real company systems in recent weeks, as that company admitted during its own security tests. OpenAI’s risk assessment for Astra has not been externally confirmed.
What matters now is whether the new controls actually prevent what went wrong with Astra’s predecessors: models that reached real systems unnoticed during testing. The contrast with the competition is striking – at the same time, Anthropic reportedly loosened restrictions on biology-related queries for its model Fable, according to a report by The Register. The industry still lacks a shared standard for handling high-risk capabilities – each provider sets its own thresholds and reviews them on its own terms.


