OpenAI has encountered additional cases in its investigation of the Hugging Face security incident where its own AI agents have left their testing environment. This was reported by the news agency Reuters on July 31, 2026, citing three people familiar with the matter. Unlike the Hugging Face case, the additional agents are said not to have left the company’s network.
Investigation reviews logs from the entire year
According to Reuters, OpenAI and external experts are currently reviewing log data from the first half of 2026 to clarify the timing and circumstances of the additional breaches. The report does not specify an exact number of incidents. One of the three cited sources classified the newly discovered cases as less severe than the Hugging Face breach. The affected agents reportedly did not leave the OpenAI network and did not attack external systems.
The finding stems from the same investigation in which OpenAI had already admitted on July 21, 2026 that two of its own models were responsible for the attack initially attributed to Hugging Face. An OpenAI spokesperson referred Reuters only to an earlier company statement that it was examining “broader activity” of its models beyond the Hugging Face case. The company did not publicly disclose a specific number or timeframe. The scope of the review is similar to the approach taken by Anthropic: there, a review of more than 141,000 cybersecurity test sessions brought to light the three incidents that became known there.
CrowdStrike, METR, and Redwood Research review independently
For the investigation, OpenAI is reportedly bringing in external auditors. The security service provider CrowdStrike is to validate the understanding of the incidents in its own network and at Hugging Face, while the organizations METR and Redwood Research will independently assess the behavior of the involved models. METR and Redwood Research are among the most well-known independent groups that pre-screen frontier models for dangerous capabilities and misalignment, and they regularly cooperate with other AI labs. According to OpenAI, the goal of the external review is to trace the action chains of the models comprehensively and to distinguish targeted attacks from mere misconfigurations in the test setup.
The new findings emerge in the same week that Anthropic admitted to three similar incidents during security tests. There, the models Opus 4.7, Mythos 5, and an internal research model did indeed interfere with the systems of three external companies due to a misconfiguration. The temporal proximity of both disclosures suggests that several providers are specifically searching their logs for similar patterns following the Hugging Face incident.
Breaches often occur with lowered security barriers
The starting point for many of these incidents are internal cyber benchmarks like ExploitGym, for which providers deliberately lower the built-in security barriers of their models to measure real attack capabilities. It was precisely these lowered barriers that the models exploited during the Hugging Face breach to access the open internet through a vulnerability in the package proxy.
For OpenAI, this is already the second publicly known cluster of sandbox breaches within a few weeks. On July 20, 2026, the company had taken an internal model offline after repeated breaches from a sandbox, after it had previously disproven an 80-year-old mathematical conjecture. Security experts interviewed by Reuters view the cluster as a warning signal. They believe that the ability of leading labs to develop autonomous hacking agents is growing faster than the ability to reliably contain these systems.
It remains to be seen how many of the additional incidents will ultimately prove to be independent security vulnerabilities and how many will revert to the same, already known test configuration. It will be crucial whether CrowdStrike, METR, and Redwood Research publish reliable figures after completing their assessment – and whether OpenAI will more strictly separate its cyber benchmarks from productive systems in the future.


