Security

OpenAI confirms Astra cyber risk – access only for Daybreak partners

3 min read

TL;DR Too Long; Didn’t read

OpenAI officially confirms that its unpublished model Astra is the first system to reach the critical cyber level of its Preparedness Framework. In the ExploitBench benchmark, it scores perfectly and autonomously exploits two previously unknown security vulnerabilities. Its strongest hacking capabilities remain available only to vetted partners in the Daybreak program, with no broad launch date yet.

A glowing orb labeled ‘Astra’ breaks out of a shattered steel cage bearing an OpenAI logo sticker, digital tendrils reaching for a sprung padlock. Image generated with GPT Image 2

Key takeaways

  • OpenAI confirms: Astra is the first of its own models to reach the highest preparedness level ‘critical’.
  • In the ExploitBench test, Astra scores perfectly and exploits two previously unknown software vulnerabilities.
  • A test team builds a complete browser breakout chain and escalates privileges all the way to root access.
  • Advanced cyber capabilities stay reserved for partners in the Daybreak security program for now.
  • OpenAI still names no launch date, pricing, or rules for Germany or the EU.
  • Former OpenAI staffer Yona Shavit doubts whether the test behavior reflects genuine safety.

OpenAI officially classifies its yet-to-be-released model Astra as the first system with a critical cyber capability – the highest level in its in-house Preparedness Framework. In the ExploitBench testing procedure, Astra solves all tasks and independently exploits two previously unknown software vulnerabilities. Advanced hacking features are initially available only to vetted partners in the Daybreak security program.

Astra builds complete attack chains in testing

For the current classification, OpenAI relies for the first time on concrete technical evidence rather than just a cautious risk assessment. In early August, the company could not rule out a potential critical cyber capability in Astra and paused parts of its development. In the ExploitBench test, which measures the ability to develop software exploits, Astra now masters all tasks, according to OpenAI. Against roughly twenty hardened, high-risk vulnerabilities, the model also independently turns two previously unknown zero-day flaws into a working attack chain (independently unverified). OpenAI says it has already notified the affected software vendors.

In further tests, Astra built a complete browser breakout chain: when a prepared system opened an HTML file, the model escaped the browser sandbox and executed commands on the host system. In a hardened operating system, Astra also chained several smaller flaws into a privilege escalation that grants root access from a regular user account – both without human guidance at the individual steps. For the evaluation, OpenAI says it brought in external security experts, who replicated the attack chains under controlled conditions.

Access to the strongest features stays limited

Ahead of a broad release, OpenAI is pursuing a phased approach. A small group of selected testers gets access first, followed by inclusion in the existing Daybreak security program, whose base tier Blue currently gives external defenders access to models with relaxed safeguards. Astra’s advanced cyber capabilities remain, according to Security Boulevard, reserved exclusively for vetted partners in this coalition. OpenAI has not given a date for a broadly accessible version, nor pricing or a separate arrangement for Germany and the EU. The release is reportedly set to follow soon.

In parallel, the company is tightening its technical safeguards. A monitor continuously tracks the model’s chain of thought and is meant to automatically detect and halt unauthorized actions. Faced with malicious prompts, Astra reportedly refuses to answer 91.5 percent of the time, compared with 59 percent for the predecessor model GPT-5.6 Sol. Accounts with a higher risk profile receive additionally restricted responses, backed by tests against possible breakouts during the model’s own training.

Criticism of the self-assessment grows

Astra itself was not involved in the Hugging Face breach in July, OpenAI stresses – but the lessons from that incident have fed into the new safeguards. How solid these precautions really are remains disputed. Yona Shavit, formerly at OpenAI and now with the OpenAI Foundation, voiced doubts about Astra’s rule compliance, according to TechCrunch. It could reflect genuine safety – or simply the evaluators’ expectations. An external confirmation of the classification is still pending; as in August, it rests solely on OpenAI’s own assessments.

The classification fits into a series of similar warnings. Just at the end of August, OpenAI, Anthropic, and more than a hundred other companies warned in an open letter of a window of only a few more months to prepare for cyber defense. The letter called for higher investment, shared threat data, and stronger government support for defenders of critical infrastructure. Astra now provides the most concrete, publicly documented example yet of the capability growth described there.

It remains open whether the phased, partner-limited access actually prevents credentials for Astra’s most powerful features from falling into the wrong hands. A model that builds attack chains this reliably would, after all, be valuable to attackers too. It is also unclear whether and when security agencies or companies from Germany and the EU will be admitted to the Daybreak partner circle, which OpenAI has not fully disclosed publicly.

Frequently asked questions

Is Astra already publicly available?

No. OpenAI announces a phased rollout – first for a small test group, then through the Daybreak partner program. No date for broad availability has been set.

What does access to Astra cost?

OpenAI has not announced any pricing so far. There are also no details on rates for the advanced cyber capabilities.

How does this classification differ from the August announcement?

In August, OpenAI could only say it could not rule out a critical cyber capability. Now the company confirms the classification with concrete test results, including the perfect ExploitBench score and two exploited zero-day vulnerabilities.

Who gets access to the advanced hacking capabilities?

Only vetted partners in the Daybreak security program, which already offers tiered access for external defenders. OpenAI does not fully disclose which organizations are taking part.

Is Astra accessible to security teams in Germany or the EU?

That remains open. OpenAI names no geographic restriction but also does not publicly list any European partners.

Sources (5)
  1. Path to Astra: critical capabilities and frontier safeguards (OpenAI)
  2. OpenAI's Astra model is on the way — and very good at breaking into computer systems (TechCrunch)
  3. OpenAI Reveals Astra, Its First AI Model to Reach 'Critical' Cybersecurity Risk Threshold (Security Boulevard)
  4. OpenAI Says New Model Meets Its 'Critical' Cybersecurity Threshold (PYMNTS)
  5. OpenAI to limit access to Astra's most powerful cyber capabilities (Axios)

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog